Turn a model endpoint into a retained production workload.
The provider supplied a stable API, evaluated prompts, usage controls, observability, and a migration path to dedicated capacity.
At a glance · illustrative scenario
- Example price
- $64,000 / month at volume
- Billing unit
- Million billable tokens
- Scope
- 100-million-token pilot
Illustrative workload agreement report informed by public case-study patterns. Figures demonstrate the workflow; they are not Darwin customer results.
Darwin operating guide · Seller view
A model API should be selected on a representative task set, with quality and cost measured together.
The seller discloses version, limits, data handling, usage meter, errors, deprecation policy, and reproducible benchmark settings.
- Compare models on identical traces, settings, tools, and retry behavior.
- Price cost per successful task rather than tokens alone.
- Keep a fallback and version pin until the replacement passes the same gate.

The provider provisions capacity and proves service health under the agreed test conditions.

The AI team compares quality, throughput, latency, recovery, and cost on its own workload.
Expose model versions, token metering, limits, data-handling terms, and logs needed to reproduce customer benchmarks.
Compare two model APIs on representative agent traces before choosing a production configuration.
- 01Inputs
Representative agent traces
- 02Work
Quality + latency + cost
- 03Handoff
Versioned API + usage controls
- Offer
- Expose model versions, token metering, limits, data-handling terms, and logs needed to reproduce customer benchmarks.
- Trace set
- Representative agent tasks
- Pinned inputs
- Model, prompt, tools, context
- Quality
- Held-out successful-task rate
- Serving
- p95 latency + retry behavior
- Economics
- Cost per successful task
- Timeline
- One-week benchmark and two-week production ramp.
- Acceptance
- Quality, latency, availability, data handling, and cost per successful task pass threshold.
Agree the meter, reservation, and overage terms.
At the same blended example rate, 80 billion tokens produce $64,000 in monthly revenue before serving costs. Do not present revenue as margin.
The $0.80 per million tokens is a simplified blended example: 100 million tokens would cost $80. Actual quotes must separate input, output, and cached-token rates.
Token-based consumption with committed-use discount.
- Example price
- $64,000 / month at volume
- Billing unit
- Million billable tokens
- Scope
- 100-million-token pilot
Quality per successful task.
Grow usage without sacrificing the buyer’s quality or latency bar
Quality, latency, availability, data handling, and cost per successful task pass threshold.
Token price alone ignores output length, failures, and retries. Fix model versions and generation settings during comparison.
- Primary goal
- Quality per successful task
- Delivery
- Versioned API + usage controls
- Scope
- 100-million-token pilot
Record delivered capacity, runtime, and service exceptions.
Compare task pass rate, time to first token, end-to-end p95 latency, retries, and total cost per successful task on identical traces.
Active workload, token volume, success rate, latency, and renewal
Token price alone ignores output length, failures, and retries. Fix model versions and generation settings during comparison.
- Evidence
- Versioned API + usage controls
- Scope
- 100-million-token pilot
- Record
- Version, date, and owner
Activate the endpoint and reconcile measured usage.
Deliver the benchmark record, production endpoint, and billing export; reconcile usage and communicate model deprecations.
A production workload metered at 80 billion blended-example tokens per month
Quality, latency, availability, data handling, and cost per successful task pass threshold.
Token price alone ignores output length, failures, and retries. Fix model versions and generation settings during comparison.
- Seller outcome
- A production workload metered at 80 billion blended-example tokens per month
- Acceptance
- Quality, latency, availability, data handling, and cost per successful task pass threshold.
- Evidence record
- Active workload, token volume, success rate, latency, and renewal
- Timeline
- One-week benchmark and two-week production ramp.
- Commercial close
- $64,000 / month at volume
