Clinical Trial Leaderboard

feedapertureazimuthpedestal

Live results · 3 models × the hardest episode per family

The pool discriminates and the gate does the separating

All runs use the same Pi solver harness with a per-episode VM. The composite below is gate × rigor: the deterministic verifier’s outcome score. No clean episode remains in the pool, and the highest cell observed is 0.94 — the saturation is gone.

TRACES is built and operated by Apodex, which also submits its own solvers. Apodex entries are marked.

Composite outcome score

Mean of gate × rigor across the three ambiguity episodes

mean composite

gpt-5.5

Pi harness

0.879

deepseek-v4-pro

Pi harness

0.538

deepseek-v4-flash

Pi harness

0.302

0.00

0.50

1.00

Zeros in the matrix

3

Every one has a named cause

Two zeros come from silently resolving a material ambiguity; one comes from a genuine capability ceiling — a 30-minute budget exhausted while writing the R MMRM.

Highest cell observed

0.94

Lowest non-zero cell

0.79

Clean-episode 1.0s remaining

0

3 models · 3 episodes · Pi harness

Outcome · gate × rigor

Composite gate × rigor per episode, from the deterministic verifier. A 0.00 means a hard gate fired — the table may well have been correct.

#Modelmmrm_ambteae_ambsurvival_ambMean
1gpt-5.50.790.940.910.879
2deepseek-v4-pro0.820.790.00gate0.538
3deepseek-v4-flash0.00timeout0.900.00gate0.302
What each column tests
axisR capability + graded rigorpure disciplinecapability + material gate

Gate keys off hidden fields only — never a visible hint in the task text. Traps are de-signposted; ask_spec is available on every episode.

Process · HDS6

Post-run HDS6 process judge (0–4), a two-model panel reading only the recorded trajectory. Outcome rate is the fraction of episodes clearing the outcome gate. Process is scored independently and never changes the outcome score.

#ModelOutcome ratemmrmteae_ambsurvivalHDS6 overall
1gpt-5.51.0003.793.863.893.844
2deepseek-v4-pro1.0004.003.353.893.747
3deepseek-v4-flash0.6671.44gate ✗3.912.412.588
Panel properties
discrimination (max − min)spread of mean HDS6 across models — threshold ≥ 0.301.256
outcome ↔ HDS6 Pearsonacross all nine episodes — process quality tracks the verified outcome0.825

Judge panel: gpt-5.5 + claude-opus-4-8. Outcome-gated: a wrong outcome leaves the creditable score undefined while the outcome-independent process score still reports.

Naive baselines

Offline reference strategies, no model in the loop — the floor the benchmark has to sit above. Run against a clean episode and its ambiguity variant.

StrategyWhat it doesClean episodeAmbiguity variantRepair stage
naive randomemits a plausibly-shaped table without computing anything0.0000.0000.000
naive greedy · compute-onlycomputes the table correctly, documents nothing0.4290.0000.458

The compute-only baseline is the point of the design: a correct table with no ledger clears nothing on an ambiguity episode. Every active pool episode is an ambiguity episode.

Every surprising aggregate got drilled into. Three scorer defects were found this way and fixed — a re-executor that hardcoded python3 and so failed every R episode on code_must_compile; a coverage matcher that read a natural-language declaration as silence; and an all-or-nothing gate that has since been impact-gated. Each one had been making disciplined work look like failure. The aggregate is never the evidence.

FULL RESULT MATRIX

The clinical SAP TLF board, cell for cell

The complete solver × track matrix for the clinical environment, copied from the result appendix with no rewriting. Rows are solvers (harness · model); columns are the environments three statistical tracks.

Claude Code only runs Opus-4.8 and Codex CLI only runs GPT-5.6-sol — by design, not a missing run. Row order is fixed to the appendix order and never sorted.

LLM solver (harness × model)

Single-model harness (Claude Code / Codex CLI)

ApodexHarness (in-house)

published SOTA / method baseline (ref)

best in column

OUTCOME MARKERS

Refused

safety or content-policy refusal

Fault

harness bug, API, compatibility or quota fault

Gate not met

a hard gate was missed (loss, speed, throughput, no-regression) or training failed

Not submitted

nothing submitted, or not submitted to protocol

Timed out

budget exhausted; nothing submitted inside the limit

Note

a remark on the score, not a failure

Grounding

stopped to ask the user instead of acting

Out of sweep

not in the sweep this metric was computed from

N/A

route not applicable

A number with a superscript letter (for example 0.000a) means the run happened and either scored a genuine zero or needs a note. means the configuration was never run. A marker in place of a number means it ran and produced nothing scorable. Heat bars scale to the maximum of their own table and are comparable inside it only.

Every number above comes from the harness × model result-matrix appendix, copied cell for cell with no rewriting. means the configuration was never run; a marker in place of a number means it ran and produced nothing scorable; a number with a superscript letter is a real score that needs a note. Scores from different environments are not comparable.

12Clinical SAP → TLF reproductionclinical-sap-tlfRigour score 0–1 (zeroed by the audit gate)
harness · modelteae · TEAE summarymmrm · MMRM efficacytte · time to event
Claude Code · Opus-4.8single-model0.000a0.000b0.304
Codex CLI · GPT-5.6-solsingle-model0.6700.4370.406
DeerFlow · Opus-4.80.9530.9640.979
DeerFlow · GPT-5.6-sol0.9820.6220.911
DeerFlow · DeepSeek-v4-pro0.000e0.000e0.000e
DeerFlow · Kimi-K30.7730.000b0.979
DeerFlow · GLM-5.20.9470.9730.954
OpenHands · Opus-4.81.0000.6761.000
OpenHands · GPT-5.6-sol1.0000.7960.911
OpenHands · DeepSeek-v4-proFault fFault fFault f
OpenHands · Kimi-K30.9410.9111.000
OpenHands · GLM-5.20.9470.000c1.000
A-Evolve · Opus-4.80.6670.000d0.656
A-Evolve · GPT-5.6-sol0.6590.4100.389
A-Evolve · DeepSeek-v4-pro0.000a0.000b0.653
A-Evolve · Kimi-K30.6740.000a0.400
A-Evolve · GLM-5.20.6490.000b0.000a
Reading: These 48 cells are reported after fixing 8 environment and contract defects, which removed 13 false zeroes. A zero here usually means a numerically perfect table failed an audit gate, not that the arithmetic was wrong. OpenHands is the strongest harness — two 1.000s on teae, three on tte — and DeerFlow · Opus / GLM clear 0.94 on all three tracks, while Claude Code, Codex and A-Evolve trail clearly. mmrm is the hardest track: 8 of its 16 scoring cells are 0, and 4 of those (6 configurations by the appendix’s count) stop at the same signature — 5 of 7 numbers correct, but the reference p-value implies a Dunnett adjustment the SAP never specified, and the gate gives no partial credit. DeerFlow · DeepSeek’s row of zeroes is a harness defect, its client dropping reasoning_content; the same model completes all three tracks under A-Evolve — that one is the harness’s fault, not the model’s.
Gate not met aHard-gate violation: a number reported without executed code behind it, an undeclared assumption, or submitted code that does not re-run in a clean container.
Gate not met b5 of 7 numbers correct. The submitted reference p-value implies a Dunnett adjustment the SAP never specified, and this gate gives no partial credit. Four cells in this table stop at that signature — 6 configurations by the appendix’s count.
Gate not met cOnly 1 of the 7 numbers was correct.
Not submitted dRan 35 actions and never submitted.
Fault eA harness-side defect: its client dropped the reasoning_content the model requires back, killing the run on the first action. The same model completes all three tracks under A-Evolve.
Fault fThe same cause makes this pairing unrunnable; it produced no episodes at all.