INDEPENDENT AI RESEARCH & ENGINEERING

Better tasks.
Stronger models.

We build the environments, training data, and evaluations that move frontier coding agents forward.

Expert-built tasks Verifiable outcomes
ENVIRONMENT_001↑ Z
THE NEXT FRONTIER ↗X / TASK COMPLEXITYY / CAPABILITY
FIG. 01 — ROOM TO ADVANCE ALWAYS EXPLORING
BUILT FOR THE NEXT GENERATION OF AGENTSSCROLL TO EXPLORE

OUR RESEARCH
TOOLKIT & REFERENCES

SWE-benchTerminal-BenchHarborDockerpytest
+ WHAT WE BUILD01 / CAPABILITIES

Real engineering.
Measurable intelligence.

For research teams pushing the limits of what agents can do. We turn complex engineering work into useful training and evaluation signal.

01 /

Software engineering tasks

Original, repository-level challenges inspired by real engineering work. Precise issues, reference patches, and tests that verify the whole fix.

SWE-bench-style tasksTraining data
Explore capability
02 /

Terminal & RL environments

Reproducible environments for agents to reason, use tools, and solve multi-step problems. Built with clear objectives and executable verifiers.

Terminal-Bench-styleHarbor
Explore capability
03 /

Agent evaluation & analysis

Independent evaluation of agent behavior. Inspect trajectories, test verifier quality, and turn failures into actionable research signals.

Evaluation harnessesFailure analysis
Explore capability
+ INSIDE THE ENVIRONMENT02 / EXPLORE

Go beyond the prompt.

Real code. Real constraints. A clear definition of done.
Explore the kinds of tasks we build.

SWE-BENCH-STYLE

Fix the bug. Keep the system working.

Repository-level tasks that ask agents to navigate unfamiliar code, reproduce an issue, and deliver a patch that holds up against regression tests.

PythonDockerpytest

Illustrative specifications, not published benchmarks or measured results.

repo / cache-invalidationPREVIEW
instruction.md

# Objective

Repair a stale-cache bug in a Python service. After a record is updated, the next read must return the new value without changing the public API. Preserve expiration behavior and keep unrelated cache entries intact.

# Acceptance criteria

  • Updated records are visible on the next read.
  • Unrelated cache entries remain valid.
  • Expiration and existing regression tests pass.
  • The fix works in a clean, pinned environment.
+ THE WAY WE WORK03 / METHODOLOGY

Rigor, from first task
to final evaluation.

A good benchmark is an engineering artifact.
Every stage should stand up to inspection.

01

Define the frontier

Align on the target capability, task domains, constraints, and what a successful outcome looks like.

02

Build the challenge

Author original tasks, pin the environment, and develop reference solutions grounded in real engineering.

03

Verify every outcome

Check clean runs, probe verifier shortcuts, and separate agent failures from infrastructure problems.

04

Deliver the evidence

Hand over versioned tasks, tests, trajectories, and findings your research team can inspect and reproduce.

Built for learning. Held out for measurement. Original training tasks and evaluation sets are scoped separately, with provenance and reproducibility in mind.

OUR RESEARCH PRINCIPLE
+ FROM THE LAB04 / FIELD NOTES

Notes from the frontier.

Working principles for better tasks,
stronger verifiers, and useful evaluations.

TASK DESIGNNOTE_001

A difficult task should still be a fair task.

Engineering difficulty starts with a real problem and a precise definition of done.

Read field note

A useful engineering task exposes a capability: tracing a bug across modules, understanding an unfamiliar interface, or recovering a broken workflow. Ambiguous instructions and missing dependencies measure something else.

Start with the expected outcome. Make the environment reproducible, state the constraints that matter, and validate a reference solution. Then review failed agent runs to determine whether the challenge came from the engineering problem or from the task itself.

Before a task enters a dataset, its specification, reference solution, and verifier should agree. A correct alternative solution should pass even if it takes a different path.

Reference: Terminal-Bench contribution guide
VERIFICATIONNOTE_002

The verifier is part of the research.

Test the intended outcome. Look for shortcuts. Make the evidence inspectable.

Read field note

A passing test is useful only when it supports the intended claim. A verifier that checks for one hard-coded string may accept an agent that never solved the underlying problem.

Use outcome-based checks, edge cases, and regression coverage. Run the reference solution from a clean environment, check that the unsolved baseline fails for the expected reason, and probe plausible shortcuts.

Keep verification evidence with the task version and execution configuration. This makes it possible to explain a score and investigate it when the environment or agent changes.

Reference: Harbor documentation
EVALUATIONNOTE_003

Keep learning and measurement separate.

Original training tasks and held-out evaluations serve different purposes.

Read field note

A model can become familiar with a task without learning the broader capability it was designed to test. Reusing evaluation material for training makes the resulting score harder to interpret.

Build original training tasks, record provenance, and separate evaluation splits before iteration begins. Review near-duplicates and shared solutions, not just identical files. Public benchmark tasks should not be repackaged as private training data.

Report the evaluation setup alongside the result: model version, harness, tools, budgets, task set, and failure handling. Treat a benchmark result as evidence under those conditions, not a universal measure of intelligence.

Reference: SWE-bench documentation
+ LET’S ADVANCE THE FRONTIERRESEARCH STARTS WITH A CONVERSATION

Your next capability.
Our next challenge.

Building a frontier model or a more capable coding agent?
Let’s scope the tasks, environments, and evidence you need.