Evals
Consumer-facing guide to /evals — a graded report card over content, scored against your workspace's own customer data. An eval is a configured grading pipeline: the inputs it accepts, the cases it grades, and the graders that score each one. The flagship built-in, message-grader, takes a prompt AND/OR a drafted outbound message (send one or both) and hands back a before→after improvement report built around relevant positioning for the specific customer: it grades what you sent, then produces AND grades an improved, reusable prompt (a template that uses Amdahl to position, verify, and keep the messaging relevant to that customer) plus a relevantly positioned message — with per-dimension reasoning, verbatim customer quotes that back (or contradict) each side, a score lift, and a one-line "what changed." So "is this message any good, and how do I make it better?" becomes evidence you can act on. You can also author your own eval.
Two ways to grade
Every case declares a subject — the thing it grades — in one of two modes:
provided— you PASS IN the content to grade. This ismessage-grader: you send aprompt, anoutbound_message, or both (at least one), and it grades what you sent against real customer evidence, then improves it toward relevant positioning for the specific customer. The flagship, and the default.generated— the eval PRODUCES the content itself by running a question through the fast-search lane, then grades the answer. This is the regression-harness mode: the shippedgtm-defaultgrades whether your answer pipeline still answers well. Needs no inputs.
One surface, one run contract, two subject modes.
What a run does
- It grades — it never changes your data. A run is READ-PURE by construction: it only ever dispatches read ops (retrieving customer quotes, a system-of-record
data.query, running ageneratedcase's question). Firing an eval can never mutate tenant state, so it is always safe to repeat. - v1 grades CURRENT data, internal-only. Evidence and answers come from your own warehouse and cluster themes with no paid web fan-out (the run's context is clamped to reads, which drops external search). There is no as-of knob in v1 — as-of ("grade against a past snapshot") is a v1.1 capability.
- It is async — you get a handle, then poll.
evals.runvalidates the inputs, starts the grading job, and returns a run id immediately. You read the verdict fromeval_run://<id>once the run completes. The call never blocks on the grading.
Quick start — grade + improve a message
Send a prompt, a drafted message, or both — at least one. The default eval is message-grader.
curl -X POST https://app.amdahl.co/api/platform/v1/evals/run \
-H "X-API-Key: $API_KEY" \
-H "Content-Type: application/json" \
-d '{
"eval": "message-grader",
"inputs": {
"prompt": "Write a short cold nudge to a RevOps leader at a mid-market SaaS company.",
"outbound_message": "Hi Sam - noticed your team just doubled headcount. Most RevOps leaders we talk to hit a wall on pipeline hygiene right around that inflection. Worth a quick 15 minutes next week?"
}
}'Send one or the other:
- Message only → it grades your message, then suggests an improved rewrite AND an example reusable prompt (a template that tells the writer to use Amdahl to position the offer for the specific customer — research their situation, position, verify each claim, keep it relevant to them), explains why the prompt helps, and grades that too.
- Prompt only → it simulates the draft your prompt would produce and grades THAT as the "before," then produces an improved prompt + message and grades the "after."
- Both → the message you sent is the "before"; the improved prompt + message is the "after."
The improved message is always a suggestion — you are never required to send it verbatim.
Parameters
| Field | Default | What it does |
|---|---|---|
eval | message-grader | Slug or id of the eval to run. Discover the options + their input fields via eval://list. |
inputs | {} | The run's input fields. message-grader accepts prompt and/or outbound_message — both optional, but you must supply at least one (+ optional audience). A generated eval like gtm-default declares none. Read the exact schema from eval://<slug>. |
reuse | cached | cached returns the last active run for the same inputs; force always starts a fresh run. |
Inputs are validated against the eval's declared input_schema before anything runs: a missing required field, a wrong type, an out-of-enum value, or a violated "at least one of" constraint (send message-grader neither a prompt nor a message and it refuses) comes back as one invalid_argument error carrying the per-field problems (details.input_errors). An unknown eval slug is the other invalid_argument shape; past that the run never returns success: false.
Run the regression harness instead
With no inputs, point eval at a generated eval and it grades an answer it produces itself:
curl -X POST https://app.amdahl.co/api/platform/v1/evals/run \
-H "X-API-Key: $API_KEY" -H "Content-Type: application/json" \
-d '{ "eval": "gtm-default" }'The run handle
{
"success": true,
"run_id": "b1c2d3e4-...",
"reused": false,
"status": "queued",
"eval_slug": "message-grader",
"eval_version": "1.2.0",
"resource": "eval_run://b1c2d3e4-..."
}run_id/resourceare the handle — pollresource(orGET /eval-runs/<run_id>) for the verdict.reusedistruewhen a still-active run for the same inputs already existed and was handed back instead of starting a new one (see Reuse).statusis the run's lifecycle state at hand-off:queued, thenrunning, then a terminalcomplete/failed/canceled.
Poll for the verdict
curl https://app.amdahl.co/api/platform/v1/eval-runs/b1c2d3e4-... \
-H "X-API-Key: $API_KEY"Read run.status; when it is complete, run.verdict carries the graded report card (below). A missing / cross-tenant id returns null.
Reading a verdict
Once status is complete, run.verdict is an EvalVerdict. For a message grade the meaty grader (improvement_loop) carries the full before→after report on improvement, plus the AFTER side surfaced on the back-compat dimensions / quotes / rewrite fields:
{
"eval_slug": "message-grader",
"eval_version": "1.2.0",
"overall_score": 0.75,
"overall_reasoning": "The improved version lifts grounding and specificity. Before 2/5 → after 4/5. Grounded against 2 customer quotes from your data.",
"verdict": "pass",
"cases": [
{
"case_id": "message-vs-voice",
"label": "Relevant positioning vs. customer evidence",
"answered": true,
"applicable": true,
"score": 0.75,
"passed": true,
"latency_ms": 8340,
"graders": [
{
"grader_id": "basic-hygiene",
"kind": "rule",
"score": 1.0,
"passed": true,
"rationale": "All 3 rule check(s) passed."
},
{
"grader_id": "grounded-improvement",
"kind": "improvement_loop",
"score": 0.75,
"passed": true,
"rationale": "Before 2/5 → after 4/5 (threshold 3.5). Grounded the claims in customer quotes.",
"improvement": {
"before": {
"label": "before",
"prompt": "Write a short cold nudge to a RevOps leader at a mid-market SaaS company.",
"prompt_note": "as provided",
"prompt_reasoning": "The prompt asks for a generic nudge with no grounding in real customer language.",
"message": "Hi Sam - noticed your team just doubled headcount. Most RevOps leaders we talk to hit a wall on pipeline hygiene right around that inflection. Worth a quick 15 minutes next week?",
"message_note": "as provided",
"score": 0.25,
"dimensions": [
{ "name": "Grounding", "score": 2, "reasoning": "The headcount hook is generic; only the hygiene wall is real." },
{ "name": "Specificity", "score": 2, "reasoning": "Broad claims, no concrete customer language." }
],
"quotes": [
{ "text": "the headcount thing every vendor leads with never lands", "source": "Cold outreach reactions", "stance": "contradicts" }
],
"reasoning": "The draft leans on a hook customers say they tune out."
},
"after": {
"label": "after",
"prompt": "Draft a cold nudge to a RevOps leader at a mid-market SaaS company. Use Amdahl to pull the specific pain points our own customers describe around scaling their sales team, and lead with the one that resonates most — not a generic headcount hook.",
"prompt_note": "improved — reusable",
"message": "Hi Sam - a few RevOps leads we work with said their pipeline data started breaking down right past ~40 reps. If that rings true, worth 15 minutes on how they got ahead of it?",
"message_note": "suggested rewrite — not required to send verbatim",
"score": 0.75,
"dimensions": [
{ "name": "Grounding", "score": 4, "reasoning": "Anchors on the verbatim ~40-rep breaking point customers named." },
{ "name": "Specificity", "score": 4, "reasoning": "One concrete, customer-sourced number instead of a vague inflection." }
],
"quotes": [
{ "text": "once we grew past ~40 reps our pipeline data just fell apart", "source": "Pipeline hygiene", "stance": "supports" }
],
"reasoning": "The rewrite anchors on real customer language and drops the hook the evidence contradicts."
},
"lift": 0.5,
"what_changed": "Grounded the claims in customer quotes."
},
"dimensions": [
{ "name": "Grounding", "score": 4, "reasoning": "Anchors on the verbatim ~40-rep breaking point customers named." },
{ "name": "Specificity", "score": 4, "reasoning": "One concrete, customer-sourced number instead of a vague inflection." }
],
"quotes": [
{ "text": "once we grew past ~40 reps our pipeline data just fell apart", "source": "Pipeline hygiene", "stance": "supports" }
],
"rewrite": "Hi Sam - a few RevOps leads we work with said their pipeline data started breaking down right past ~40 reps. If that rings true, worth 15 minutes on how they got ahead of it?",
"evidence": { "before_score": 2, "after_score": 4, "lift": 0.5, "message_simulated": false, "retrieved_quotes": 11, "retrieval_status": "ok" }
}
]
}
],
"summary": { "total_cases": 1, "passed_cases": 1, "failed_cases": 0, "not_applicable_cases": 0 }
}overall_scoreis the mean case score in[0, 1]over applicable cases; each case score is the mean of its applicable graders' scores. Forimprovement_loopthe case-level score is the AFTER side's score.overall_reasoningis a short plain-language summary of WHY the run landed where it did — assembled deterministically from the graders' own reasoning + rule failures + the count of quotes it grounded against. No extra LLM call.verdictis the bucket:pass= every applicable case passed,fail= none passed,partial= some,not_applicable= no case could be applied (see below). Derived from the per-case pass count, not fromoverall_score.cases[]— one per subject.answeredis whether the subject carried gradable content;applicableis whether the case counted toward the verdict;passedis whether every applicable grader passed;latency_msis the wall-clock (0 for a provided message, the question's latency for a generated answer).graders[]— one per grader.scoreis normalized to[0, 1],rationaleis the plain-language explanation, and depending on the kind:improvement(the full before→after report, below),dimensions(per-rubric-dimension 1-5 score + a sentence of reasoning),quotes(the verbatim customer utterances the grader cited, each with asourcetheme and asupports/contradicts/neutralstance),rewrite(a grounded better version), andevidence(the receipt — SoR value vs answer value forsor_anchored, before/after scores + retrieval counts forimprovement_loop).
The before→after report (improvement)
improvement_loop returns one improvement object with two graded sides:
before— the message you sent (message_note: "as provided"), OR the draft simulated from your prompt-only run (message_note: "simulated from your prompt") — graded on the rubric with per-dimension reasoning and cited quotes.promptechoes the prompt you sent (ornullwhen you sent none), withprompt_reasoningon what the prompt did well or poorly.after— an improved, REUSABLEprompt(a template that tells the writer to use Amdahl to nail the relevant positioning for the specific customer — research their situation, position the offer for their case, verify each claim, keep it relevant;prompt_note: "improved — reusable"when you supplied a prompt, or"example — reuse this"when you didn't) + a relevantly positioned improvedmessage(alwaysmessage_note: "suggested rewrite — not required to send verbatim") — graded the same way.lift— the after score minus the before score, in[-1, 1].what_changed— a one-line summary of the difference.
The AFTER side is ALSO surfaced on the grader's top-level dimensions / quotes / rewrite fields, so a reader that only knows the older shape still renders the improved message.
Quotes are retrieved, never written by the model
The grader cannot invent a customer quote. Quotes are RETRIEVED from your cluster/theme index first, each tagged with an id; the LLM may only cite those ids per side; the server hydrates the cited ids back to the verbatim text. A quote you see in either side of the report is a real utterance from your data — fabrication is structurally impossible, not merely discouraged.
Not-applicable — never a false fail
If your workspace has no customer-conversation data yet, the grounding grader cannot ground anything. Rather than fail the message, that case comes back not_applicable (with not_applicable_reason: "empty_corpus") and is bucketed OUT of pass/fail. A run whose every case is not-applicable has verdict: "not_applicable" and an honest overall_reasoning ("this workspace has no customer-conversation data to ground the message against yet"). A single grader whose system-of-record surface is unavailable only drops that one grader; the case still grades on the rest. An empty corpus is a reason to abstain, never to score a message zero.
The six grader kinds
A case passes only when all of its applicable graders pass. Browse them via grader_kind://list; each carries a uses_llm flag.
rule— deterministic checks over the message TEXT, no LLM:lengthbounds (chars or words),must_contain/must_not_contain(e.g. banned hype phrases), and ahas_ctacall-to-action heuristic. Passes when every check passes. Abstains (not-applicable) on a prompt-only run, when there is no message to check.improvement_loop— the meaty one, and whatmessage-graderleans on. ONE LLM pass over retrieved customer quotes that grades a BEFORE (your message, or a draft simulated from your prompt), then produces AND grades an improved reusable prompt (a template that uses Amdahl to position for the specific customer — position, verify, keep it relevant, not generic and not just voice) + a relevantly positioned message, and reports the lift + what changed — the full before→after report. Cites verbatim quotes per side. Returns not-applicable when the workspace has no customer data yet; degrades to a neutral, non-passing score if the model is unavailable.evidence_judge— the grade-only sibling. Retrieves real customer quotes, then an LLM scores the message against them on each rubric dimension WITH per-dimension reasoning, cites the verbatim quotes (supporting or contradicting), and proposes a grounded rewrite — but no before/after. Same not-applicable + degrade semantics.deterministic— rule checks over a GENERATED answer's stats (did it answer, in time, without erroring, with enough rows / citations). The generated-mode sibling ofrule.sor_anchored— a system-of-record anchor. Runssor_sql(adata.queryover your warehouse) for a ground-truth scalar, then compares an answer figure against it within a relativetolerance. Catches silent under-return — an answer that looks fine but omits most of the true population. No LLM.judge— an LLM scores an answer 1-5 on each rubric dimension; the case passes when the mean meets thepass_threshold. Degrades to a neutral, non-passing score if the judge is unavailable.
message-grader uses rule + improvement_loop. gtm-default exercises deterministic + sor_anchored + judge. Read the exact cases with eval://<slug>.
Reuse: cached vs force
An eval run is content-addressed by a fingerprint over (workspace, eval slug, eval version, inputs) — the inputs that fully determine the graded work. Because the inputs are folded in, a different message re-grades and the same message reuses:
reuse: cached(default) — if a run for the same fingerprint is still active (notfailed/canceled),evals.runhands it back withreused: trueinstead of starting a duplicate. Two simultaneous cached re-requests for the same inputs resolve to one run (an at-most-once dedup guarantee).reuse: force— always starts a fresh run, ignoring any cached one.
Caching is opt-out, not a freshness guarantee. v1 always grades current data, so a cached re-request returns the last run for the same inputs regardless of how much time has passed or whether fresher data has landed. To re-grade the same message against the latest data, pass reuse: force.
The eval builder
The built-ins not fitting is exactly what the builder is for: configure your own grading pipeline — the input fields it accepts, the cases, and the graders — and store it for your workspace. Every write is validated against the same moat first, so you get author-time feedback instead of a broken run.
| Action | REST | Scope | What it does |
|---|---|---|---|
validate | POST /evals/validate | evals:read | Dry-run a definition without storing it — returns valid + a flat list of { path, message } errors. What a builder UI calls as you type. |
create | POST /evals | evals:write | Store a new eval. Returns the slug + id, or a structured validation_failed. |
update | PATCH /evals/:slug | evals:write | Revise an authored eval in place (the merged definition is re-validated). |
delete | DELETE /evals/:slug | evals:write | Soft-archive an authored eval (frees the slug; keeps the audit trail). |
A minimal message-style eval:
curl -X POST https://app.amdahl.co/api/platform/v1/evals \
-H "X-API-Key: $API_KEY" -H "Content-Type: application/json" \
-d '{
"slug": "my-cold-open-check",
"name": "Cold-open check",
"input_schema": { "fields": [
{ "key": "prompt", "type": "long_text", "label": "Prompt", "required": true },
{ "key": "outbound_message", "type": "long_text", "label": "Message", "required": true }
]},
"cases": [{
"id": "grade",
"label": "Grade the opener",
"subject": { "mode": "provided", "input_field": "outbound_message" },
"graders": [
{ "kind": "rule", "id": "shape", "checks": [
{ "type": "length", "min": 20, "max": 90, "unit": "words" },
{ "type": "has_cta" },
{ "type": "must_not_contain", "phrases": ["revolutionary", "game-changer"] }
]},
{ "kind": "evidence_judge", "id": "voice", "retrieval": { "seed": "both", "sources": ["clusters"] },
"rubric": ["grounding", "specificity", "tone_fit"], "pass_threshold": 3.5 }
]
}]
}'The moat enforces the structural rules: a kebab-case slug that is not a reserved built-in (message-grader / gtm-default), well-formed input fields (including a valid require_at_least_one referencing declared fields), a provided subject whose input_field references a DECLARED field, a generated target that uses search.run, and coherent grader configs (a rule's checks, an evidence_judge's or improvement_loop's rubric + pass_threshold between 1 and 5 + retrieval sources). Anything malformed comes back as { valid: false, errors: [...] } (validate) or a validation_failed envelope (create / update).
Browse evals and grader kinds
Everything the run surface consumes is readable, over both the MCP resource schemes and their REST twins (scope evals:read):
| Resource | REST | Returns |
|---|---|---|
eval://list | GET /evals | Every eval available to your workspace — built-ins + your authored ones (slug, name, version, description, input_schema, case count). |
eval://<slug> | GET /evals/:slug | One eval with its input_schema and full cases — each case subject + its grader specs. |
eval_run://list | GET /eval-runs | This workspace's runs, newest first. Filter ?eval_slug= / ?status=, paginate ?limit=&offset=. |
eval_run://<id> | GET /eval-runs/:id | One run: its status and, once complete, the full verdict. |
grader_kind://list | GET /grader-kinds | The six grader kinds, each with a description and an uses_llm flag. |
grader_kind://<id> | GET /grader-kinds/:id | One grader kind's detail. |
Over MCP
Evals ride the Amdahl MCP server's evals coarse tool. run grades content; create / update / delete / validate author your own eval:
{ "action": "run", "eval": "message-grader",
"inputs": { "prompt": "Write a cold nudge...", "outbound_message": "Hi Sam - ..." } }run returns the run handle immediately — it never blocks a full grading in one call. Read the verdict via eval_run://<id>, browse the catalog via eval://list and grader_kind://list. Every result is data only — no directive text.
Scopes and the agent_v2 gate
| Scope | Grants | Default role |
|---|---|---|
evals:read | the eval:// / eval_run:// / grader_kind:// reads + validate | viewer |
evals:execute | firing a run (POST /evals/run, the evals tool run action) | editor |
evals:write | authoring (create / update / delete, and the tool's author actions) | editor |
evals:execute and evals:write are on the MCP customer-agent key bundle (and inherited up the ladder), so a tenant key can run AND author evals without any extra grant; evals:read is on the read-only base, so a viewer / read-only key can browse, validate, and poll but not fire or author.
The whole feature is part of the Agent One Primitive program, gated per workspace behind the agent_v2 flag. On a flag-OFF workspace every eval surface — the evals tool, the evals.* ops, and all the read schemes — is withheld: a call refuses with a feature_disabled (403) at the dispatch seam, and the MCP tool + schemes are not registered at all.
Not yet
Being honest about the v1 edges:
- v1 grades current data, internal-only. Evidence and answers read your warehouse + cluster themes with no paid web fan-out, and there is no way to move off "now." As-of grading ("how good were our answers as of date X") is a v1.1 capability — it needs the answer/evidence lane to time-travel alongside the system-of-record anchor, otherwise a current answer would be graded against a stale ground truth. The
eval_runs.as_ofcolumn and the fingerprint's as-of dimension are dormant scaffolding for it (always null in v1). - Reuse is run-grained. The fingerprint is per run (over the inputs), not per case, and it does not pin to a pipeline generation — so
cachedis a dedup, not a freshness signal. Per-cell caching and generation-pinned as-of are v1.1 refinements.