Deliver a benchmark the buyer can rerun after every change.
Experts built realistic tasks, calibrated scoring criteria, and passed independent review before the suite became a reusable release control.
At a glance · illustrative scenario
- Example price
- $320 / accepted task
- Billing unit
- Accepted task
- Scope
- 120-task suite / 20-task batch
Illustrative evaluation report informed by public case-study patterns. Figures demonstrate the workflow; they are not Darwin customer results.
Darwin operating guide · Seller view
A release suite should recreate the work, not ask generic model questions.
A seller is delivering reproducible task worlds, expected behavior, criterion-level rubrics, and resolved QA—not a favorable model score.
- Keep task sources and expected outcomes out of the candidate system’s context.
- Freeze the suite before comparing models, prompts, tools, or serving configurations.
- Report repeated-run consistency and critical failures beside the aggregate score.

Specialists score real work against calibrated rubrics and document disagreements.

The product team reviews repeat-run findings, failed traces, and the release gate.
Author a 20-task batch with source material, expected behavior, scoring rubrics, and reproducible environments.
Compare two finance-agent configurations on 120 held-out month-end tasks spanning source files, calculations, reconciliations, and review notes.
- 01Inputs
Support policy + task sources
- 02Work
Rubrics + independent review
- 03Handoff
Versioned suite + run results
- Offer
- Author a 20-task batch with source material, expected behavior, scoring rubrics, and reproducible environments.
- Task worlds
- 8 reproducible companies
- Held-out suite
- 120 workflow tasks
- Run design
- 5 attempts per task
- Scoring
- Criterion rubrics + adjudication
- Gate
- Critical failures block release
- Timeline
- Six weeks from workflow sampling to release review.
- Acceptance
- Every task is reproducible, source-grounded, independently reviewed, and mapped to a release criterion.
Agree rates for accepted expert work.
At $320 per accepted task, a 20-task batch invoices $6,400. Agree rejection reasons, a correction window, and payment after batch acceptance.
The $48,000 example includes $38,400 for 120 accepted tasks at $320 each, plus $9,600 for calibration, runs, and independent QA.
Fixed scope with payment tied to independently accepted tasks.
- Example price
- $320 / accepted task
- Billing unit
- Accepted task
- Scope
- 120-task suite / 20-task batch
Release readiness.
Produce defensible tasks and rubrics for a live agent workflow
Every task is reproducible, source-grounded, independently reviewed, and mapped to a release criterion.
A high aggregate score cannot override a critical failure. Passing this suite does not establish performance outside its coverage.
- Primary goal
- Release readiness
- Delivery
- Versioned suite + run results
- Scope
- 120-task suite / 20-task batch
Make every judgment reproducible and reviewable.
Log rubric pass rate by failure category, reviewer agreement, runtime, and cost per successful task. Freeze the suite before comparing configurations.
Acceptance rate, reviewer agreement, and unresolved ambiguity
A high aggregate score cannot override a critical failure. Passing this suite does not establish performance outside its coverage.
- Evidence
- Versioned suite + run results
- Scope
- 120-task suite / 20-task batch
- Record
- Version, date, and owner
Deliver a reusable suite and settle accepted tasks.
Hand over the accepted 20 tasks and revision log, then invoice the approved batch. A model failing the tasks is not a failure of the author’s delivery.
20 accepted tasks with source packages and scoring guides
Every task is reproducible, source-grounded, independently reviewed, and mapped to a release criterion.
A high aggregate score cannot override a critical failure. Passing this suite does not establish performance outside its coverage.
- Seller outcome
- 20 accepted tasks with source packages and scoring guides
- Acceptance
- Every task is reproducible, source-grounded, independently reviewed, and mapped to a release criterion.
- Evidence record
- Acceptance rate, reviewer agreement, and unresolved ambiguity
- Timeline
- Six weeks from workflow sampling to release review.
- Commercial close
- $320 / accepted task
