Skip to content

How Darwin works
for evaluation specialists.

A practical guide to turning expert judgment into accepted tasks, calibrated reviews, traceable evidence, and repeat evaluation work.

Evaluation specialists can package domain judgment as a reliable quality system—not a pile of disconnected scores.

For specialists supplying the quality layer

  • Domain experts
  • Code evaluators
  • Red-team firms
  • QA leads
  • Evaluation studios
  • Rubric designers
  • Annotation teams
  • Safety researchers
  • Benchmark authors
  • Review operations teams

Prove domain coverage on the buyer’s task family with a paid sample, expected output, rubric draft, reviewer method, and declared conflicts.

01 · Provider qualificationSeller brief

Prove expertise against the actual task family.

The provider is qualified on the buyer’s domain and output format before receiving held-out authoring work. Coverage and conflicts are visible at assignment time.

Domain
Accounting operations
Qualification
Source-grounded sample + gold output
Experience
Credentialed practitioners
Batch
20 accepted tasks
Conflicts
Declared before assignment

Darwin evaluation record · illustrative workflow

Deliver a benchmark the buyer can rerun after every change.

Experts built realistic tasks, calibrated scoring criteria, and passed independent review before the suite became a reusable release control.

At a glance · illustrative scenario

Example price
$320 / accepted task
Billing unit
Accepted task
Scope
120-task suite / 20-task batch

Illustrative evaluation report informed by public case-study patterns. Figures demonstrate the workflow; they are not Darwin customer results.

Darwin operating guide · Seller view

A release suite should recreate the work, not ask generic model questions.

A seller is delivering reproducible task worlds, expected behavior, criterion-level rubrics, and resolved QA—not a favorable model score.

  • Keep task sources and expected outcomes out of the candidate system’s context.
  • Freeze the suite before comparing models, prompts, tools, or serving configurations.
  • Report repeated-run consistency and critical failures beside the aggregate score.
Release suite · both sides of the exchangeExpert judgment becomes a release decision.
Two accounting evaluators independently checking an AI reconciliation against source evidence
Seller perspective

Specialists score real work against calibrated rubrics and document disagreements.

An AI product and finance team tracing a failed reconciliation before release
Buyer perspective

The product team reviews repeat-run findings, failed traces, and the release gate.

Author a 20-task batch with source material, expected behavior, scoring rubrics, and reproducible environments.

Compare two finance-agent configurations on 120 held-out month-end tasks spanning source files, calculations, reconciliations, and review notes.

How this evaluation worksIllustrative workflow
  1. 01Inputs

    Support policy + task sources

  2. 02Work

    Rubrics + independent review

  3. 03Handoff

    Versioned suite + run results

Offer
Author a 20-task batch with source material, expected behavior, scoring rubrics, and reproducible environments.
Task worlds
8 reproducible companies
Held-out suite
120 workflow tasks
Run design
5 attempts per task
Scoring
Criterion rubrics + adjudication
Gate
Critical failures block release
Timeline
Six weeks from workflow sampling to release review.
Acceptance
Every task is reproducible, source-grounded, independently reviewed, and mapped to a release criterion.

Agree rates for accepted expert work.

At $320 per accepted task, a 20-task batch invoices $6,400. Agree rejection reasons, a correction window, and payment after batch acceptance.

The $48,000 example includes $38,400 for 120 accepted tasks at $320 each, plus $9,600 for calibration, runs, and independent QA.

Fixed scope with payment tied to independently accepted tasks.

Example price
$320 / accepted task
Billing unit
Accepted task
Scope
120-task suite / 20-task batch

Release readiness.

Produce defensible tasks and rubrics for a live agent workflow

Every task is reproducible, source-grounded, independently reviewed, and mapped to a release criterion.

A high aggregate score cannot override a critical failure. Passing this suite does not establish performance outside its coverage.

Primary goal
Release readiness
Delivery
Versioned suite + run results
Scope
120-task suite / 20-task batch

Make every judgment reproducible and reviewable.

Log rubric pass rate by failure category, reviewer agreement, runtime, and cost per successful task. Freeze the suite before comparing configurations.

Acceptance rate, reviewer agreement, and unresolved ambiguity

A high aggregate score cannot override a critical failure. Passing this suite does not establish performance outside its coverage.

Evidence
Versioned suite + run results
Scope
120-task suite / 20-task batch
Record
Version, date, and owner

Deliver a reusable suite and settle accepted tasks.

Hand over the accepted 20 tasks and revision log, then invoice the approved batch. A model failing the tasks is not a failure of the author’s delivery.

20 accepted tasks with source packages and scoring guides

Every task is reproducible, source-grounded, independently reviewed, and mapped to a release criterion.

A high aggregate score cannot override a critical failure. Passing this suite does not establish performance outside its coverage.

Seller outcome
20 accepted tasks with source packages and scoring guides
Acceptance
Every task is reproducible, source-grounded, independently reviewed, and mapped to a release criterion.
Evidence record
Acceptance rate, reviewer agreement, and unresolved ambiguity
Timeline
Six weeks from workflow sampling to release review.
Commercial close
$320 / accepted task

Blog

On Agentic Commerce, Part I: Intent

Why intent—not checkout—is the defining interface of agentic commerce, and why turning goals into outcomes is a network problem.

Company · Sanjit Juneja
Read article

$1M to build the supply network for agentic commerce

Company · Darwin
Read article
Explore the blog