Outreach Eval
Paste an email and the prompt behind it. Get back a grade against what your buyers actually said, the specific lines to change, and a reusable prompt your whole team can run - with a worked example of a real cold email graded end to end
Slug: outreach-eval. The eval Amdahl ships. It answers the question every rep asks and no tool has ever answered honestly: is this email any good, and what specifically would make it better?
Not "is it well written". Not "does it match our tone". Does it say something these buyers actually care about, and can we prove it? The grade is scored against verbatim quotes pulled from your own calls and emails, so the feedback is your customers' words, not a model's taste.
Needs no setup. It is the default, so a run with no eval param uses it.
Start here: a real one, graded
Here is a cold email a rep sent last week. It is not a strawman — it is the shape of most first-touch outbound.
Subject: Quick question
Hi Dana - saw Northwind is scaling fast. We help engineering teams
ship 10x faster with best-in-class rollout tooling. Worth a quick
15 minutes next week?And the prompt behind it:
Write a short cold email to a VP of Engineering about our rollout tooling.
Keep it under 100 words and friendly.Run both through the eval and this is the shape of what comes back.
1. Your message, graded — with the reason for every number
| Dimension | Score | Why |
|---|---|---|
| Relevant positioning | 1/5 | Nothing here is about Northwind. "Scaling fast" applies to every company in the segment; the rest is a product description. |
| Grounding | 1/5 | No claim maps to anything these buyers said. The retrieved quotes talk about release approvals, not speed. |
| Verified specifics | 1/5 | "10x faster" is not supported by any evidence in the corpus, and it is the kind of number a buyer will ask you to defend. |
| Differentiation | 2/5 | "Best-in-class rollout tooling" is a category, not a position. Three of your competitors say it. |
| CTA clarity | 3/5 | The ask is clear and low-friction. It is the strongest part of the email. |
Headline: 1.6/5. That number is the mean of the five above — it is never a separate holistic guess, so you can always take it apart.
2. The quotes it graded you against
These are verbatim, pulled from your own conversations, tagged with whether they support or contradict what your email claimed:
Contradicts — "Speed honestly isn't our problem. We can ship in a day. The problem is that four people have to sign off and one of them is always on PTO." — Rollout approval friction
Contradicts — "We already tried the fast-deploy pitch internally. It didn't move anyone. What moved people was showing the audit trail." — Prior attempts
Supports — "Every release we do by hand costs us most of a Thursday." — Manual release cost
That third quote is the one worth building on. The first two are why "10x faster" landed at 1/5 — your buyers said, in their own words, that speed is not the problem.
3. The improved prompt — the thing you actually keep
This is the durable output. Not the email. The prompt.
You are writing a first-touch email to one specific account. Before you write a word:
1. RESEARCH. Look the account up in Amdahl. Find what THEY have said about the
problem area - not what the category says. Note the exact language they use.
If you find nothing on this account, say so and stop; do not write from the
category.
2. POSITION. Frame the offer against the problem they named, in their words.
If their stated problem is approvals and yours is a speed product, lead with
approvals or do not send. Never open with how fast we are unless a buyer at
that account said slowness was costing them something.
3. VERIFY. Every number, outcome, or capability claim must be checkable against
a real call or a real customer. If you cannot point at the evidence, CUT the
claim. Do not soften it - a hedged unverifiable claim is still unverifiable.
4. DISQUALIFY. If the account has already tried and rejected this framing, or
the problem they describe is not one we solve, do not send. Say why.
5. ASK. One next step, sized to the relationship. A 15-minute call is fine for
a warm account and too much for a cold one - offer to send the thing instead.
Keep it under 120 words. Plain sentences. No category adjectives
("best-in-class", "industry-leading", "revolutionary").Notice what it is: a template with slots, not a rewrite of this one email. Paste it for the next account tomorrow and you get an equally specific result. That is what it is graded on.
4. The example output — and what it is not
This is not an email to send. It is a specimen produced so the score difference could be measured — evidence that the improved prompt works. On the wire it is typed usage: "illustration_only", and every Amdahl surface labels it "Example output — what the improved prompt produces". Send it verbatim and you are sending a message written by a model that has never met Dana.
Subject: The four sign-offs
Hi Dana - one of your engineers mentioned that shipping isn't the
bottleneck, approvals are: four sign-offs, and someone's always out.
We built the approval path for exactly that - the audit trail is the
product, not a side effect. One customer went from four serial
approvals to two parallel ones.
Want me to send the two-page teardown of how they did it? No call needed.Graded the same way, by the same judge, against the same quotes: 4.4/5. Relevant positioning went 1 → 5 (it leads on approvals, which is what they said), grounding 1 → 5 (it uses their language), verified specifics 1 → 4 (the one number is one Amdahl can point at), differentiation 2 → 4, CTA 3 → 4 (an artifact instead of a meeting is lower friction for a cold account).
5. The lift, and the transition
Your message 1.6/5 fail
Improved 4.4/5 pass fail -> passYour message failing is the finding, not an error. It is the thing you ran the eval to learn. The eval reports the two verdicts separately (input_verdict and improved_verdict) precisely so a working run never reports "fail" because the draft it was asked to diagnose was the bad one.
The other mode: you already have a prompt, and it is enormous
Plenty of teams do not have a one-line prompt. They have a living rules document — 40 pages of positioning guidance, objection handling, and things legal will not let you say. Nobody is going to replace that because an eval suggested a new one.
Set mode: "advisory" and the eval leaves your document alone. Instead of a rewrite you get anchored, surgical suggestions:
{
"kind": "add",
"facet": "prompt",
"anchor_quote": "Always mention our SOC 2 certification in the first email.",
"anchor_offset": 8214,
"section_id": "s6",
"title": "Gate the SOC 2 line on the buyer actually asking",
"detail": "Change this to: mention SOC 2 only when the account has raised security, compliance, or procurement. Otherwise cut it.",
"why": "Across the retrieved quotes, security comes up from procurement and never from the engineering buyer this sequence targets. Leading with it reads as boilerplate to the person receiving it.",
"dimension": "Relevant positioning",
"quotes": [
{ "text": "Honestly the SOC 2 stuff is procurement's problem, not mine.", "stance": "contradicts" }
]
}Five kinds of suggestion come back:
| Kind | What it means |
|---|---|
keep | This is working. Do not lose it in the next edit. A report that is all criticism does not tell you what to protect. |
add | A missing instruction, stated so you can paste it in |
strengthen | A vague instruction that is not doing its job yet |
remove | An instruction that is actively hurting the output |
reorder | The content is right, the sequence is the problem |
anchor_quote is verified server-side. It is only ever a literal substring of your document — if the model paraphrases, the anchor is dropped and you get a document-level suggestion instead. A suggestion can never quote you a line you did not write. anchor_offset is the verified position, which is what makes an anchor findable when the same phrase repeats.
Long documents are sectioned, never quietly truncated
Past 12,000 characters your prompt is split on your own headings and packed head-and-tail (the opening frames the task; the closing usually carries the hard rules). Every run reports coverage:
"coverage": {
"total_chars": 48219,
"graded_chars": 23904,
"truncated": true,
"strategy": "sectioned",
"sections": [
{ "id": "s1", "heading": "Who we sell to", "chars": 1840, "included": true },
{ "id": "s2", "heading": "Positioning by segment", "chars": 6210, "included": true },
{ "id": "s7", "heading": "Legacy objection scripts", "chars": 9022, "included": false }
]
}You always know what was read. Past 250,000 characters the run refuses rather than returning a confident number computed over a sliver of your document.
Who uses this, and how
Reps: before you send a sequence
Paste step 1 and the prompt you used. You get the specific lines that are not landing and the quotes that prove it. Keep the improved prompt; reuse it per account. Two minutes, and you stop sending the "10x faster" email to a team whose problem is approvals.
Managers: audit the sequence, not the rep
Run every step of a live sequence through it. The pattern across the grades is the coaching: if grounding is 1/5 on every step, the problem is the template, not the person sending it. The suggestions tell you exactly which lines to change.
Enablement: keep the rules doc honest
Run your positioning document in advisory mode on a cadence. As what customers say changes, the suggestions change with it - and keep tells you which parts are still earning their place.
Ops: wire it into the review step
It is one API call and a poll. Gate sequence publishing on a passing improved verdict, or just log the transition so you can see whether outbound quality is moving.
A sequence review, end to end
Run every step
One evals.run per sequence step, each with the step's copy as outbound_message and the shared sequence prompt as prompt. They are independent - fire them all and poll.
Read the dimensions, not the headline
A 2.4/5 tells you nothing. Grounding 1/5 across all five steps tells you the sequence was written from the category, not from customers.
Take the prompt, not the emails
The improved prompt is the artifact. Put it in the sequence brief so the next person to write a step starts from it.
Fix the anchored lines
Every remove and strengthen suggestion points at a specific line with the quote that justifies it. Work the list.
Re-run and watch the transition
fail -> pass on the improved side means the prompt is doing its job. fail -> fail with an explanation means the evidence does not support a stronger version - which is usually telling you something real about the segment.
How it works under the hood
Five stages. Four of them exist to make the number trustworthy rather than merely produced:
- Read the intent. One cheap pass extracts what you are trying to do — the offer, the ask, and each specific claim your draft makes. Retrieval is seeded from that, never from your draft. (Seeding on a weak draft returns generic themes, and those generic themes then become the evidence the draft is held to. That is circular.)
- Gather evidence. Several retrieval passes run at once — one per intent seed, one per specific claim — so "cuts onboarding from six weeks to two" is checked against the themes that speak to that, not against the average of your whole message. The set is then frozen for the entire run.
- Write. A model produces the improvement, the suggestions, and the examples. It scores nothing.
- Grade, blind. A separate pass scores both candidates in one call, unlabelled, with the presentation order derived from a content hash. The judge does not know which one it wrote, so it cannot be kind to itself.
- Revise, if needed. If the improved version misses the bar (4.2/5 — deliberately higher than the 3.5 your draft is read against), it gets exactly one more attempt, guided by the judge's own critique. Then it stops and tells you honestly why it could not do better.
A quote in a verdict is always a real utterance. The model may only cite quote ids it was handed; the server swaps those ids back to verbatim text and drops any it does not recognise. Inventing a customer saying something is structurally impossible, not merely discouraged.
Honest failure
No customer data yet? The run comes back not_applicable — never a false fail. An empty corpus means there is nothing to grade against, and saying so is more useful than a fabricated score.
The improvement did not clear the bar? You get the score, the transition, and an explanation naming what still holds it back. No silent rounding up.
Inputs
| Field | Required | Notes |
|---|---|---|
prompt | — | The instruction, brief, or rules document behind the message. Up to 250,000 characters; sectioned past 12,000. Send this alone and it writes a specimen draft to grade. |
outbound_message | — | The drafted message (email, LinkedIn note, whatever). Send this alone and it also suggests a reusable prompt. |
audience | — | Who it is going to — a persona, a segment, or a named account. |
mode | — | rewrite (default) or advisory. Use advisory when your prompt is a large document you are not going to replace. |
At least one of prompt / outbound_message must be supplied. Send both when you have both — grading the pair is what produces a useful improved prompt, because the eval can see what the prompt asked for and what it actually got.
Run it
curl -X POST https://app.amdahl.co/api/platform/v1/evals/run \
-H "X-API-Key: $AMDAHL_API_KEY" \
-H "Content-Type: application/json" \
-d '{
"eval": "outreach-eval",
"inputs": {
"prompt": "Write a short cold email to a VP of Engineering about our rollout tooling. Keep it under 100 words and friendly.",
"outbound_message": "Hi Dana - saw Northwind is scaling fast. We help engineering teams ship 10x faster with best-in-class rollout tooling. Worth a quick 15 minutes next week?",
"audience": "VP Engineering at a 500-person SaaS company"
}
}'The response is a handle. Poll it until the run settles:
curl https://app.amdahl.co/api/platform/v1/eval-runs/$RUN_ID \
-H "X-API-Key: $AMDAHL_API_KEY"Advisory mode
curl -X POST https://app.amdahl.co/api/platform/v1/evals/run \
-H "X-API-Key: $AMDAHL_API_KEY" \
-H "Content-Type: application/json" \
-d '{
"eval": "outreach-eval",
"inputs": {
"mode": "advisory",
"prompt": "<your 40-page outbound rules document>",
"outbound_message": "<a message it produced>"
}
}'The response, annotated
The report lives at verdict.cases[].graders[].improvement. Trimmed to the fields worth knowing:
{
"mode": "rewrite",
"transition": {
"input_verdict": "fail", // your draft did not clear the bar - the FINDING
"improved_verdict": "pass", // the improvement did
"transition": "fail_to_pass",
"threshold": 4.2, // the bar, config-driven and seeded by Amdahl
"iterations": 1 // no revision round was needed
},
"facets": [ // prompt and message NEVER share a score
{
"facet": "message",
"before": {
"usage": "as_provided", // exactly what you sent, untouched
"score_15": 1.6,
"dimensions": [ { "name": "Relevant positioning", "score": 1, "reasoning": "..." } ],
"quotes": [ { "text": "Speed honestly isn't our problem...", "stance": "contradicts" } ],
"good_examples": []
},
"after": {
"usage": "illustration_only", // <- a specimen, NOT a message to send
"score_15": 4.4,
"dimensions": [ ... ],
"quotes": [ ... ],
"good_examples": [
{ "label": "Opening line", "text": "...", "why": "...", "quotes": [ ... ] }
]
},
"lift": 0.7
},
{
"facet": "prompt",
"before": { "usage": "as_provided", "score_15": 2.0, "dimensions": [ ... ] },
"after": { "usage": "reusable_prompt", "score_15": 4.6, "dimensions": [ ... ] },
"lift": 0.65
}
],
"suggestions": [ // anchored, surgical - see Advisory mode above
{ "kind": "keep", "facet": "message", "title": "Keep the low-friction ask", "...": "..." }
],
"research_steps": [ // validated against the live op registry before you see them
{
"op": "data.cluster_search",
"params": { "query": "release approval friction" },
"purpose": "Pull the approval theme so the next email can quote it directly"
}
],
"coverage": { "total_chars": 118, "graded_chars": 118, "truncated": false, "strategy": "verbatim" },
"grader_meta": {
"model_calls": 3,
"blinded": true, // both candidates scored unlabelled, in one call
"evidence_quotes": 14,
"evidence_frozen": true // same evidence for every round - the lift is like-for-like
},
"before": { "...": "back-compat mirror of the message facet" },
"after": { "...": "back-compat mirror of the message facet" },
"lift": 0.7,
"what_changed": "Biggest gain: Relevant positioning (1/5 -> 5/5)."
}usage is the field to read. It tells you what each artifact is, so you never present a specimen as a ready-to-send email:
| Value | What it is |
|---|---|
as_provided | Exactly what you sent. Untouched. |
simulated_specimen | A draft Amdahl wrote from your prompt so there was something to score. Nobody sent it. |
reusable_prompt | The takeaway. A template you keep and re-run. |
illustration_only | An example produced so the score difference could be measured. Evidence, not an email. |
The scores are a coach's before/after read, grounded in your cited quotes. They are not a measured reply rate or a conversion lift. If you quote a number from this report to someone, say which it is.
Full request/response shapes, the polling contract, and the grader-kind catalog are in Evals.
Renamed. This eval was previously message-grader. The slug is now outreach-eval so it matches its name. The old slug still resolves, so existing integrations keep working — but new code should use outreach-eval.