Skip to content

feat(evals): add LLM-as-judge scoring #8514

Description

@sudoKrishna

Problem

The eval suite scores answers with substrings and regexes. Those measure
phrasing, not correctness: across three live runs every failure was a valid
paraphrase or an over-specific assertion, not a wrong answer. Open-ended
behavior — is the answer grounded in the tool results? does it answer the
question? — cannot be scored this way at all.

Proposal

Add LLM-as-judge scoring alongside the deterministic checks.

  • A judge model scores an answer against a weighted rubric (grounding,
    completeness, …) and returns structured numbers.
  • runScenario gains an optional judge; when present it adds a judge check.
  • The judge transport is an injectable OpenAI-compatible completion, so a
    recorded transcript can replay it deterministically in CI.

Location: apps/sim/evals/agent-tool-use/judge.ts.

Acceptance criteria

  • judgeAnswer scores a rubric and returns a weighted verdict
  • Malformed/missing-score judge responses are rejected, scores clamped
  • Key-free unit tests for parsing, weighting, and clamping
  • A live test proves a grounded answer outscores an invented one
  • runScenario can attach a judge check

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    featureNew feature or request

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions