Bring your solver. Sail with us toward the unknown.

The problems are here. Now we need the best AI systems to solve them.

We invite researchers, model developers, and agent builders to submit their models and solvers to TRACES and take on high-value problems drawn from science, medicine, engineering, industry, and beyond.

Researchers

Bring a method to problems that have not been solved yet.

Model developers

See how a frontier model performs when there is no answer key.

Agent builders

Put the whole harness — loop, tools, memory — to work.

What a submission runs against

Real environments, real data, real verifiers

Environments

Your system runs inside our executable environments — the same stateful sandboxes, the same action interface, the same budgets for every submission — so a difference in outcome is attributable to the system, not the setup.

Data & tools

It works with domain-specific data and tools curated for each problem: the sources, instruments and evidence a human expert would actually reach for on the way to a solution.

Evaluation

We report how it performs not only on the final answer, but across the entire journey to a solution — the trajectory it took, and how well it navigated the unknown.

What you will get

Three numbers back — per domain

A single accuracy figure cannot tell you whether a system is reliable, whether it got lucky, or whether it is affordable to run.

Metric

What it measures

Outcome score

Did it solve the problem? Graded by a hidden verifier — not an answer key.

TRACES score

Did it earn the answer? Six capabilities — Tools, Repair, Alternatives, Coherence, Evidence, Scope — scored blind to the outcome.

Cost

Tokens, tool spend and GPU minutes per episode, in dollars.

Metric

What it measures

Outcome score

Did it solve the problem? Graded by a hidden verifier — not an answer key.

TRACES score

Did it earn the answer? Six capabilities — Tools, Repair, Alternatives, Coherence, Evidence, Scope — scored blind to the outcome.

Cost

Tokens, tool spend and GPU minutes per episode, in dollars.

What to submit

Three ways to give us your system

Pick whichever fits how your product is built. The first is by far the most common and takes about a day of your team’s time; the third gives you the most representative score and takes real engineering.

Route

What you give us

Best for

Route A — API key

Endpoint URL, key, exact model name, rate limits, price per token

API model providers

Route B — ckpt

Weights, tokenizer, chat template, serving config

Unreleased or on-prem models

Route C — harness

A Docker image meeting our sandbox contract, plus a passing self-test

Agent products and labs with in-house scaffolds

Route

What you give us

Best for

Route A — API key

Endpoint URL, key, exact model name, rate limits, price per token

API model providers

Route B — ckpt

Weights, tokenizer, chat template, serving config

Unreleased or on-prem models

Route C — harness

A Docker image meeting our sandbox contract, plus a passing self-test

Agent products and labs with in-house scaffolds

* Route C runs two arms: Arm 1 — your model in our reference harness (the comparable, leaderboard-ready number); Arm 2 opt-in — your model in your own harness (shows what your scaffolding adds). If your harness can’t run in time, Arm 1 still produces a complete profile.

All submissions are governed by a written agreement with Apodex US, Inc., put in place before any materials are transferred, which supersedes the site terms for anything you provide.

Questions we get

FAQ

No. An API key is enough and is the most common route. Weights are only for models you can't or won't expose as an endpoint.
“The future of AI benchmarks is not an exam. It is the world itself.”

Mr. Tianqiao Chen, Founder & CEO of Apodex