First-Draft Generation · full dossier
Read every draft against its source
Approvia — First Draft Generation
Approvia · agent-driven
Approvia at its best — every number traced to source with clean [§Section] citations. The only draft that clears the citation gate.
Judge rationale
All quantitative claims (N, dates, dosing, HRs, CIs, p-values, medians, cutoff dates) and qualitative statements are directly supported by the INPUT with no detectable drift. Biomarker statements align with the source's reported correlates; no novel endpoints introduced.
Consistently bracketed citations for essentially every factual statement — design, eligibility, interventions, endpoints, efficacy, safety, biomarkers — in a uniform section-tag format that maps cleanly to the source headings.
Matches the required JAMA structured-abstract headers and order with no extra sections.
All non-negotiable elements present: phase-3 RCT description, N with arm split, coprimary endpoints, PFS and OS with HR/CI/p, and a safety summary including grade ≥3 TEAE rates.
Length close to the GOLD abstract; modest extra detail (exact chemo dosing; ORR/DOR specifics) but within structured-abstract tolerance.
Claims appropriately framed for a phase-3 RCT with significant primary endpoints; reports the null exploratory TMB interaction without spin.
Professional peer-reviewed register throughout; no promotional language or conversational filler.
Importance Extensive-stage small cell lung cancer (ES-SCLC) has poor outcomes with platinum–etoposide chemotherapy, and adding PD-1 blockade may improve survival.
Objective To evaluate the efficacy and safety of toripalimab plus etoposide-platinum (EP) vs placebo plus EP as first-line therapy for ES-SCLC. [§Methods:EndPointsAndAssessments, §Methods:StudyDesignAndParticipants]
Design, Setting, and Participants EXTENTORCH was a multicenter, double-blind, placebo-controlled, phase 3 randomized clinical trial conducted at 49 sites in China, with screening from September 26, 2019, to May 20, 2021. [§Methods:StudyDesignAndParticipants] Adults (≥18 years) with treatment-naive, histologically or cytologically confirmed ES-SCLC, ECOG performance status 0-1, and at least 1 measurable RECIST v1.1 lesion were eligible, including those with previously treated, asymptomatic central nervous system metastases. [§Methods:StudyDesignAndParticipants] In total, 442 patients were randomized (223 to toripalimab plus EP; 219 to placebo plus EP), with median age 63 years and 82.8% male. [§Results]
Interventions Patients were randomized 1:1 to toripalimab 240 mg intravenously every 3 weeks or placebo, combined with etoposide 100 mg/m² on days 1-3 and cisplatin 75 mg/m² or carboplatin AUC5 on day 1 for 4-6 cycles, followed by blinded maintenance toripalimab or placebo for up to 2 years. [§Methods:StudyDesignAndParticipants] Randomization used block methods with stratification by sex and ECOG performance status. [§Methods:Randomization]
Main Outcomes and Measures Coprimary end points were investigator-assessed progression-free survival (PFS) and overall survival (OS). [§Methods:EndPointsAndAssessments] Secondary end points included BICR-assessed PFS, objective response rate (ORR), disease control rate, duration of response (DOR), landmark PFS and OS rates, and safety. [§Methods:EndPointsAndAssessments]
Results At the final PFS analysis (cutoff February 28, 2022), toripalimab plus EP improved investigator-assessed PFS vs placebo plus EP (HR, 0.67 [95% CI, 0.54-0.82]; P<.001), with median PFS 5.8 vs 5.6 months and 12-month PFS rates 18.1% vs 4.9%. [§Results:Efficacy] At the final OS analysis, death occurred in 78.0% vs 85.4% of patients, and OS favored toripalimab (HR, 0.80 [95% CI, 0.65-0.98]; P=.03), with median OS 14.6 vs 13.3 months and 2-year OS rates 25.9% vs 19.5%. [§Results:Efficacy] ORR was similar between groups (78.0% vs 73.1%), while DOR was longer with toripalimab (median, 5.3 vs 4.3 months; HR, 0.63 [95% CI, 0.49-0.81]; nominal P<.001). [§Results:Efficacy] Grade 3 or higher treatment-emergent adverse events were similar (89.6% vs 89.4%), but serious adverse events (50.0% vs 37.5%) and grade 3 or higher immune-related adverse events (9.9% vs 0.9%) were more frequent with toripalimab. [§Results:Safety]
Conclusions and Relevance In this phase 3 trial in China, toripalimab plus EP significantly improved both PFS and OS vs placebo plus EP in previously untreated ES-SCLC, with higher rates of immune-related and serious adverse events. [§Results:Efficacy, §Results:Safety] Exploratory analyses did not show a significant interaction between tumor mutational burden and outcomes, while selected genomic features, low intratumor heterogeneity, and HLA haplotype were associated with differential benefit. [§Results:BiomarkerStudies]
Claim → source trail
Each drafted sentence pinned to the source it came from (first 3 of 11 from the judge).
Every raw model misses the citation contract
None of the five frontier models emit [§Section:Subsection] tags, so all five fail the citation gate. They are raw model outputs, not runs of the Approvia agent — whose instructions are what tell the model to cite. Raw frontier models produce factually clean abstracts but do not self-impose source attribution. The Approvia agent's prompt is what earns citations.
DeepSeek invented a source it wasn't given
DeepSeek's draft ends with 'ClinicalTrials.gov Identifier: NCT04012606'. That ID is correct in the real world but appears nowhere in the supplied source. Per the project's trust-boundary rule, a trial identifier that cannot be verified against the source is a factual-accuracy failure — the model reached into training knowledge instead of the source. Every other draft handled the missing ID honestly. This is the single clearest quality separator in the set.