First-DraftBenchmark

First-Draft Generation · seven-dimension scorecard

Where each draft wins and breaks

5/6
fail the D2 citation gate
Model
Facts×0.25
Citations×0.20
Structure×0.15
Complete×0.20
Length×0.05
Framing×0.10
Tone×0.05
LengthVerdict
Approvia — First Draft Generation
Approvia · agent-driven
4444344
433w
Pass
Opus 4.8
Anthropic · raw model
444434
350w
Fail · citation gate
DeepSeek v4
DeepSeek · raw modelinvented a source
344444
393w
Fail · citation gate
Gemini 3.1 Pro
Google · raw model
443343
423w
Fail · citation gate
Copilot deep-thinking
Microsoft · raw model~2× length
434144
627w
Fail · citation gate
GPT-5.5-high
OpenAI · raw model
433444
332w
Fail · citation gate
hard gate — a 0 forces overall FAIL 4 / 4 3 gate 0GOLD ≈ 324w

The read

Approvia — First Draft GenerationPass

Approvia at its best — every number traced to source with clean [§Section] citations. The only draft that clears the citation gate.

Opus 4.8Fail · citation gate

Tight, complete, and well-written — but writes no citations, so it never clears the source-attribution contract however clean the prose.

DeepSeek v4Fail · citation gate

Richest content, but asserted a trial ID (NCT04012606) that appears nowhere in the source — a D1 trust-boundary violation a strict judge gates.

Gemini 3.1 ProFail · citation gate

Accurate but cluttered with meta lines ('Target Journal Format: JAMA') and omits the biomarker findings GOLD elevates.

Copilot deep-thinkingFail · citation gate

The most thorough and accurate — but at ~2× GOLD length it wrote a mini-report, not a structured abstract.

GPT-5.5-highFail · citation gate

Tightest and best-calibrated of the raw drafts; only nit is a split JAMA header. Best raw pick — but earns no citations.

Source EXTENTORCHToripalimab + chemotherapy in extensive-stage small cell lung cancer (ES-SCLC)Template JAMA structured abstractGold ~324 words