Software engineering tasks
Original, repository-level challenges inspired by real engineering work. Precise issues, reference patches, and tests that verify the whole fix.
Explore capabilityINDEPENDENT AI RESEARCH & ENGINEERING
We build the environments, training data, and evaluations that move frontier coding agents forward.
For research teams pushing the limits of what agents can do. We turn complex engineering work into useful training and evaluation signal.
Original, repository-level challenges inspired by real engineering work. Precise issues, reference patches, and tests that verify the whole fix.
Explore capabilityReproducible environments for agents to reason, use tools, and solve multi-step problems. Built with clear objectives and executable verifiers.
Explore capabilityIndependent evaluation of agent behavior. Inspect trajectories, test verifier quality, and turn failures into actionable research signals.
Explore capabilityReal code. Real constraints. A clear definition of done.
Explore the kinds of tasks we build.
Repository-level tasks that ask agents to navigate unfamiliar code, reproduce an issue, and deliver a patch that holds up against regression tests.
Illustrative specifications, not published benchmarks or measured results.
# Objective
Repair a stale-cache bug in a Python service. After a record is updated, the next read must return the new value without changing the public API. Preserve expiration behavior and keep unrelated cache entries intact.
# Acceptance criteria
A good benchmark is an engineering artifact.
Every stage should stand up to inspection.
Align on the target capability, task domains, constraints, and what a successful outcome looks like.
Author original tasks, pin the environment, and develop reference solutions grounded in real engineering.
Check clean runs, probe verifier shortcuts, and separate agent failures from infrastructure problems.
Hand over versioned tasks, tests, trajectories, and findings your research team can inspect and reproduce.
Built for learning. Held out for measurement. Original training tasks and evaluation sets are scoped separately, with provenance and reproducibility in mind.
OUR RESEARCH PRINCIPLEWorking principles for better tasks,
stronger verifiers, and useful evaluations.
Engineering difficulty starts with a real problem and a precise definition of done.
Read field noteA useful engineering task exposes a capability: tracing a bug across modules, understanding an unfamiliar interface, or recovering a broken workflow. Ambiguous instructions and missing dependencies measure something else.
Start with the expected outcome. Make the environment reproducible, state the constraints that matter, and validate a reference solution. Then review failed agent runs to determine whether the challenge came from the engineering problem or from the task itself.
Before a task enters a dataset, its specification, reference solution, and verifier should agree. A correct alternative solution should pass even if it takes a different path.
Reference: Terminal-Bench contribution guideTest the intended outcome. Look for shortcuts. Make the evidence inspectable.
Read field noteA passing test is useful only when it supports the intended claim. A verifier that checks for one hard-coded string may accept an agent that never solved the underlying problem.
Use outcome-based checks, edge cases, and regression coverage. Run the reference solution from a clean environment, check that the unsolved baseline fails for the expected reason, and probe plausible shortcuts.
Keep verification evidence with the task version and execution configuration. This makes it possible to explain a score and investigate it when the environment or agent changes.
Reference: Harbor documentationOriginal training tasks and held-out evaluations serve different purposes.
Read field noteA model can become familiar with a task without learning the broader capability it was designed to test. Reusing evaluation material for training makes the resulting score harder to interpret.
Build original training tasks, record provenance, and separate evaluation splits before iteration begins. Review near-duplicates and shared solutions, not just identical files. Public benchmark tasks should not be repackaged as private training data.
Report the evaluation setup alongside the result: model version, harness, tools, budgets, task set, and failure handling. Treat a benchmark result as evidence under those conditions, not a universal measure of intelligence.
Reference: SWE-bench documentationBuilding a frontier model or a more capable coding agent?
Let’s scope the tasks, environments, and evidence you need.