FORGE 2027
Mon 26 - Tue 27 April 2027 Dublin, Ireland
co-located with ICSE 2027

Call for Paper

High-quality datasets, sound evaluation methodologies, and reliable benchmarks are essential for advancing foundation models (FMs) and FM-based systems for software engineering.

The FORGE 2027 Data and Benchmarking Track provides a dedicated forum for rigorous research on the data, benchmarks, metrics, tools, and evaluation practices needed to develop and assess these systems.

The track welcomes contributions that go beyond merely applying existing models to standard evaluation settings. Submissions should provide a substantive contribution through new or improved datasets, benchmarks, evaluation methodologies, tools, empirical findings, replications, or critical analyses of current practices.

We particularly encourage work that advances data quality, methodological validity, transparency, reproducibility, responsible evaluation, and the practical relevance of FM research for software engineering.


Scope

The track welcomes two main types of submissions: Data Papers and Benchmarking Papers.

Contributions may address any software engineering task, including code generation, repository-level development, issue resolution, code review, testing and test generation, debugging, program repair, requirements engineering, software design, documentation, migration, modernization, and maintenance. We also welcome work addressing under-represented programming languages, legacy technologies, specialized domains, and industrial software ecosystems.

A submission does not need to propose a new model or learning technique. The contribution may lie entirely in the data, benchmark, evaluation methodology, tooling, or empirical findings, and submissions will be assessed on those grounds.

Conversely, papers that only apply existing models to a new task or dataset, without a substantive contribution in either dimension, are out of scope and may be rejected without full review.

Authors of submissions contributing to both categories should select the category that best represents the paper’s primary contribution. This selection will primarily support reviewer assignment and will not restrict the topics addressed in the paper.

Data Papers

Data Papers should introduce, analyze, improve, or critically assess datasets and data-centric methods supporting the development and evaluation of FMs and FM-based systems for software engineering.

Relevant contributions include, but are not limited to:

  • New datasets or carefully curated collections derived from previously available data.
  • Multilingual, multimodal, cross-project, cross-domain, or temporally evolving datasets.
  • Data generators, simulation environments, and reinforcement-learning environments.
  • Data-centric methods and tools for measuring or improving data quality.
  • Methods for data collection, cleaning, deduplication, annotation, curation, validation, documentation, and versioning.
  • Studies of data contamination, benchmark leakage, memorization, duplication, provenance, and temporal overlap.
  • Dataset audits identifying significant limitations, biases, errors, or inconsistencies in existing datasets.
  • Frameworks and practices for responsible dataset development, governance, licensing, privacy, security, fairness, and ethical use.
  • Advanced practices in data collection and curation, including cases in which the complete dataset cannot be publicly released because of legal, ethical, privacy, contractual, security, or commercial constraints.
  • Empirical studies providing substantive new insights into how data characteristics affect model behavior, performance, robustness, or reliability.

Benchmarking Papers

Benchmarking Papers should introduce or critically study benchmarks, evaluation methodologies, metrics, tools, or infrastructures for assessing FMs and FM-based systems in software engineering.

Relevant contributions include, but are not limited to:

  • New benchmarks, evaluation protocols, metrics, or benchmarking tools.
  • Execution-based and sandboxed evaluation infrastructures, containerized evaluation harnesses, and approaches for handling flaky, non-deterministic, or environment-dependent tests.
  • Dynamic, continuously updated, temporally grounded, or contamination-resistant benchmarks.
  • Evaluation frameworks for agentic, interactive, multimodal, retrieval-augmented, or tool-using systems.
  • Methods for assessing functional correctness, robustness, reliability, security, fairness, efficiency, scalability, cost, latency, or energy consumption.
  • Evaluation of non-deterministic systems, including sampling strategies, repeated measurements, variance analysis, confidence intervals, and sensitivity to model configurations.
  • Studies of the effects of model updates, API changes, model deprecation, and version drift on the reproducibility and comparability of benchmark results.
  • Human evaluation methodologies, annotation protocols, and analyses of inter-rater agreement.
  • Studies of automated evaluators, including LLM-as-a-judge approaches, their validity, reliability, biases, and limitations.
  • Benchmark audits revealing saturation, contamination, construct-validity problems, unreliable test cases, or unintended incentives.
  • Systematic evaluations of existing systems that use novel experimental designs or datasets and produce substantive new insights.
  • Replications and reproductions of previously published benchmark results, including well-supported negative or contradictory findings.
  • Reproducible benchmarking infrastructures supporting fair and consistent comparisons.
  • Meta-evaluations examining whether existing benchmarks and metrics measure the intended software engineering capabilities.

Evaluation Criteria

Data Papers

Data Papers will be evaluated according to:

  • Relevance and significance: Importance of the contribution to the FORGE and software engineering communities.
  • Data quality and methodological rigor: Soundness of the collection, generation, annotation, cleaning, validation, and analysis processes.
  • Originality and insight: Novelty of the contribution and its ability to enable new research or understanding.
  • Transparency and reusability: Quality of the documentation concerning provenance, structure, intended uses, limitations, licensing, accessibility, and maintenance.
  • Responsible development: Appropriate consideration of privacy, consent, intellectual property, security, bias, fairness, and potential misuse, where applicable.
  • Presentation quality: Clarity, completeness, organization, and positioning with respect to related work.

Benchmarking Papers

Benchmarking Papers will be evaluated according to:

  • Relevance and significance: Importance of the benchmark or evaluation problem to the FORGE and software engineering communities.
  • Originality and insight: Novelty of the benchmark, methodology, metric, tool, infrastructure, or empirical findings.
  • Methodological and construct validity: Soundness of the experimental design, baselines, metrics, statistical analyses, and evidence that the benchmark measures the intended capabilities.
  • Reproducibility and transparency: Quality and availability of the evaluation protocols, data, code, prompts, configurations, execution environments, and documentation.
  • Potential for adoption and impact: Extent to which the results, benchmarks, tools, or datasets can be adopted, extended, or used by researchers and practitioners.
  • Presentation quality: Clarity, completeness, organization, and positioning with respect to related work.

Submission Instructions

Paper Length

Submissions may contain a maximum of 4 pages, plus 1 page for references.

Appendices included in the submitted PDF count toward the page limit.

Supplementary documentation, such as dataset cards, datasheets, schema descriptions, annotation guidelines, provenance documentation, or benchmark-harness documentation, may be included in the accompanying artifact and does not count toward the paper page limit.

The paper must nevertheless contain all information necessary to understand and assess the main contribution.

Artifacts

Authors are strongly encouraged to provide the datasets, code, prompts, configurations, evaluation scripts, execution environments, and documentation needed to assess and reproduce the contribution.

For Data Papers, the dataset or data-related artifact should normally be made available to reviewers in an anonymized form at submission time. For Benchmarking Papers introducing a new benchmark, evaluation infrastructure, or tool, the corresponding benchmark data, harness, or implementation should likewise normally be available for review.

Exceptions are permitted when sharing is prevented by legal, ethical, privacy, contractual, security, or commercial constraints. Such restrictions must be clearly explained, and authors should provide sufficient documentation, representative samples, aggregate information, or alternative validation mechanisms to allow reviewers to assess the contribution.

All submitted artifacts must comply with the double-anonymous reviewing policy. Authors are responsible for ensuring that artifact URLs, repository metadata, commit histories, account names, package names, documentation, and file metadata do not reveal their identities.

Artifact availability will be considered as part of the evaluation where it is relevant to the paper’s claims. However, submissions will not be rejected solely because an artifact cannot be publicly released when the restrictions are adequately justified and the contribution can still be assessed.

Authors of accepted papers are encouraged to release the final version of their artifacts through a stable and persistent repository whenever legally and ethically possible.

Ethics, Licensing, and Data Protection

Authors are responsible for ensuring that the data they collect, use, process, and release comply with the licenses and terms of use of the original sources and with applicable data-protection regulations.

The paper or accompanying artifact documentation should describe, where applicable, the provenance and license of the data, restrictions on access or redistribution, the handling of personal or sensitive information, intended uses and limitations, and foreseeable risks, biases, or potential misuse.

Studies involving human participants should report the relevant ethical approval or explain why such approval was not required.

Format and Reviewing

All submissions must be written in English, submitted in PDF format, and conform at the time of submission to the official ACM Primary Article Template.

LaTeX users should use the sigconf, review, and anonymous options:

\documentclass[sigconf,review,anonymous]{acmart}

Submissions will undergo double-anonymous peer review. Authors must remove names, affiliations, acknowledgements, funding information, repository ownership information, and other identifying details from the submitted paper and artifacts. Prior work by the authors should be discussed in the third person.

For further guidance, authors should consult the ICSE Research Track Q&A on double-anonymous reviewing.

Other Policies

The FORGE 2027 policies on originality and simultaneous submission, use of generative-AI tools, accessibility, confidentiality, and registration and presentation apply to this track. Authors should consult the main FORGE 2027 submission pages for the authoritative versions of these policies.

Submission Site

Submissions must be made through the track submission site:

https://forge2027-benchmarking.hotcrp.com/


Questions

Questions about the Data and Benchmarking Track may be addressed to the track co-chairs:


Important Dates

All deadlines are at 23:59 Anywhere on Earth (AoE).

  • Paper submission deadline: November 15, 2026
  • Author notification: January 4, 2027
  • Camera-ready submission: January 23, 2027