Problem
The eval suite scores answers with substrings and regexes. Those measure
phrasing, not correctness: across three live runs every failure was a valid
paraphrase or an over-specific assertion, not a wrong answer. Open-ended
behavior — is the answer grounded in the tool results? does it answer the
question? — cannot be scored this way at all.
Proposal
Add LLM-as-judge scoring alongside the deterministic checks.
- A judge model scores an answer against a weighted rubric (grounding,
completeness, …) and returns structured numbers.
runScenario gains an optional judge; when present it adds a judge check.
- The judge transport is an injectable OpenAI-compatible completion, so a
recorded transcript can replay it deterministically in CI.
Location: apps/sim/evals/agent-tool-use/judge.ts.
Acceptance criteria
Problem
The eval suite scores answers with substrings and regexes. Those measure
phrasing, not correctness: across three live runs every failure was a valid
paraphrase or an over-specific assertion, not a wrong answer. Open-ended
behavior — is the answer grounded in the tool results? does it answer the
question? — cannot be scored this way at all.
Proposal
Add LLM-as-judge scoring alongside the deterministic checks.
completeness, …) and returns structured numbers.
runScenariogains an optionaljudge; when present it adds ajudgecheck.recorded transcript can replay it deterministically in CI.
Location:
apps/sim/evals/agent-tool-use/judge.ts.Acceptance criteria
judgeAnswerscores a rubric and returns a weighted verdictrunScenariocan attach a judge check