Meridian Field Services: instructional example, not project evidence
One model role at pilot: RANKING (order the eligible technician set for a job). Blended cost per resolved task, latency at the committed percentile.
| Candidate | Rank agreement | Blended cost / task | p95 @ pilot concurrency | Verdict |
|---|---|---|---|---|
| Small-fast model | 0.87 | 0.4 ¢ | 620 ms: inside the 800 ms budget (NFR-P2[1]) | SELECTED |
| Mid-tier model | 0.89 | 2.1 ¢ | 1,340 ms: blows the budget | rejected |
Verdict rationale: two points of agreement do not buy a busted latency budget and five times the cost; logged (ranking-model decision[2]), eval run attached (eval runs[3]). Pin: the versioned model id held in platform config; the floating alias is banned from production. Drift tripwire: the nightly eval runs the pinned version and the provider's newest candidate; a two-point agreement drop opens a ticket (alerting[4]).
| Re-evaluation trigger (event) | What re-runs |
|---|---|
| Provider deprecation notice on the pinned version | Full matrix |
| Any release in the small-fast class | Candidate row |
| Agreement drift below 0.84 on the nightly | Full matrix + golden-set review |
Golden set: 500 historical assignments with the dispatcher's actual choice and the eligible set at decision time, drawn from the legacy system's seven-year history (migration inventory[5]); refreshed quarterly with pilot decisions; the set is a test asset (register[6]) cited by id, not duplicated.
Metric fit: rank agreement (the model's top three contains the dispatcher's choice), a deterministic overlap measure; no rubric needed for a closed output space. LLM-as-judge: NOT USED, as a claim: the task has ground truth (what the dispatcher chose), so a judge would add bias without adding signal.
Lever ladder: prompt engineering carried the quality bar alone (PR-01[7]); retrieval adds nothing (the eligible set is already the context); fine-tuning rejected with the ladder documented: 0.87 agreement meets the pilot bar (success measures[8]) without tuning cost or drift risk.
| Role | Primary | Fallback | Class | Trigger |
|---|---|---|---|---|
| ranking | Pinned small-fast model via the gateway (INT-4[9]) | No model: the full technician list with the degradation notice (suggestions_unavailable[10]) | DEGRADED, declared & signed; quality delta: the dispatcher ranks manually, the pre-AI baseline; signed A. Reyes | 2 s timeout · 1 in-window retry · breaker at 5 failures (cited from error recovery[11]) |
No capability-equivalent fallback exists at pilot (single provider, the documented SPOF, provider strategy[11]), stated, not hidden. Routing lives in gateway config; a chain edit is a config change, no deploy.
| Question | Owner | Answer by | Blocks |
|---|---|---|---|
| Does the quarterly golden-set refresh reweight toward pilot decisions? (The legacy set encodes pre-suggestion habits, the thing the pilot is changing.) | A. Reyes with the QA owner | Oct 10, 2026 | The eval-criteria refresh rule only |