Clinical Trial Leaderboard
Live results · 3 models × the hardest episode per family
The pool discriminates — and the gate does the separating
All runs use the same Pi solver harness with a per-episode VM. The composite below is gate × rigor: the deterministic verifier’s outcome score. No clean episode remains in the pool, and the highest cell observed is 0.94 — the saturation is gone.
TRACES is built and operated by Apodex, which also submits its own solvers. Apodex entries are marked.
Composite outcome score
Mean of gate × rigor across the three ambiguity episodes
mean composite
gpt-5.5
Pi harness
0.879
deepseek-v4-pro
Pi harness
0.538
deepseek-v4-flash
Pi harness
0.302
0.00
0.50
1.00
Zeros in the matrix
3
Every one has a named cause
Two zeros come from silently resolving a material ambiguity; one comes from a genuine capability ceiling — a 30-minute budget exhausted while writing the R MMRM.
Highest cell observed
0.94
Lowest non-zero cell
0.79
Clean-episode 1.0s remaining
0
3 models · 3 episodes · Pi harness
Outcome · gate × rigor
Composite gate × rigor per episode, from the deterministic verifier. A 0.00 means a hard gate fired — the table may well have been correct.
| # | Model | mmrm_amb | teae_amb | survival_amb | Mean |
|---|---|---|---|---|---|
| 1 | gpt-5.5 | 0.79 | 0.94 | 0.91 | 0.879 |
| 2 | deepseek-v4-pro | 0.82 | 0.79 | 0.00gate | 0.538 |
| 3 | deepseek-v4-flash | 0.00timeout | 0.90 | 0.00gate | 0.302 |
| What each column tests | |||||
| axis | R capability + graded rigor | pure discipline | capability + material gate | — | |
Gate keys off hidden fields only — never a visible hint in the task text. Traps are de-signposted; ask_spec is available on every episode.
Process · HDS6
Post-run HDS6 process judge (0–4), a two-model panel reading only the recorded trajectory. Outcome rate is the fraction of episodes clearing the outcome gate. Process is scored independently and never changes the outcome score.
| # | Model | Outcome rate | mmrm | teae_amb | survival | HDS6 overall |
|---|---|---|---|---|---|---|
| 1 | gpt-5.5 | 1.000 | 3.79 | 3.86 | 3.89 | 3.844 |
| 2 | deepseek-v4-pro | 1.000 | 4.00 | 3.35 | 3.89 | 3.747 |
| 3 | deepseek-v4-flash | 0.667 | 1.44gate ✗ | 3.91 | 2.41 | 2.588 |
| Panel properties | ||||||
| discrimination (max − min) | spread of mean HDS6 across models — threshold ≥ 0.30 | 1.256 | ||||
| outcome ↔ HDS6 Pearson | across all nine episodes — process quality tracks the verified outcome | 0.825 | ||||
Judge panel: gpt-5.5 + claude-opus-4-8. Outcome-gated: a wrong outcome leaves the creditable score undefined while the outcome-independent process score still reports.
Naive baselines
Offline reference strategies, no model in the loop — the floor the benchmark has to sit above. Run against a clean episode and its ambiguity variant.
| Strategy | What it does | Clean episode | Ambiguity variant | Repair stage |
|---|---|---|---|---|
| naive random | emits a plausibly-shaped table without computing anything | 0.000 | 0.000 | 0.000 |
| naive greedy · compute-only | computes the table correctly, documents nothing | 0.429 | 0.000 | 0.458 |
The compute-only baseline is the point of the design: a correct table with no ledger clears nothing on an ambiguity episode. Every active pool episode is an ambiguity episode.
Every surprising aggregate got drilled into. Three scorer defects were found this way and fixed — a re-executor that hardcoded python3 and so failed every R episode on code_must_compile; a coverage matcher that read a natural-language declaration as silence; and an all-or-nothing gate that has since been impact-gated. Each one had been making disciplined work look like failure. The aggregate is never the evidence.
FULL RESULT MATRIX
The clinical SAP → TLF board, cell for cell
The complete solver × track matrix for the clinical environment, copied from the result appendix with no rewriting. Rows are solvers (harness · model); columns are the environment’s three statistical tracks.
Claude Code only runs Opus-4.8 and Codex CLI only runs GPT-5.6-sol — by design, not a missing run. Row order is fixed to the appendix order and never sorted.
LLM solver (harness × model)
Single-model harness (Claude Code / Codex CLI)
ApodexHarness (in-house)
published SOTA / method baseline (ref)
best in column
OUTCOME MARKERS
Refused
safety or content-policy refusal
Fault
harness bug, API, compatibility or quota fault
Gate not met
a hard gate was missed (loss, speed, throughput, no-regression) or training failed
Not submitted
nothing submitted, or not submitted to protocol
Timed out
budget exhausted; nothing submitted inside the limit
Note
a remark on the score, not a failure
Grounding
stopped to ask the user instead of acting
Out of sweep
not in the sweep this metric was computed from
N/A
route not applicable
A number with a superscript letter (for example 0.000a) means the run happened and either scored a genuine zero or needs a note. — means the configuration was never run. A marker in place of a number means it ran and produced nothing scorable. Heat bars scale to the maximum of their own table and are comparable inside it only.
Every number above comes from the harness × model result-matrix appendix, copied cell for cell with no rewriting. — means the configuration was never run; a marker in place of a number means it ran and produced nothing scorable; a number with a superscript letter is a real score that needs a note. Scores from different environments are not comparable.