Skip to content

fix(gooddata-eval): let --gate decide a what-if item, not pass@K alone - #1831

Open
Tomkess wants to merge 1 commit into
masterfrom
fix/gate-what-if
Open

Tomkess wants to merge 1 commit into
masterfrom
fix/gate-what-if

Conversation

@Tomkess

@Tomkess Tomkess commented Sep 24, 2026

Copy link
Copy Markdown
Contributor

The bug

run_agentic_what_if computes pass_power_k. The evaluator never read it:

if not summary.pass_at_k:      # pass^K computed, then ignored

So --gate power on a flaky item passes on one good run out of K — the exact case the gate exists to catch, and the report says the opposite of what happened.

This is live: what-if merged this morning in #1799.

The fix

What-if now does what the other eight multi-run kinds already do:

  • takes gate: EvalGate = DEFAULT_GATE, and the dispatch passes it
  • asks gate_passed(...) for the verdict
  • carries gate_failure_note(...) into the assertion message — this matters, because the message body describes the best run, which under pass^K can be a run that passed, so without the note a failing item reads like a passing one
  • stamps the gate on the run metadata and publishes pass@K / pass^K / gate_passed, so a gated run is readable in Langfuse and not only in the exit code

Default behaviour is unchanged. Without --gate this is pass@K exactly as before, pinned by its own test.

Verification

Three tests: the power gate failing a 1-of-2 item, any still passing the same item, and the default matching any. Reverting the one-line verdict change makes the power test fail, so it is testing the fix rather than the scaffolding.

1209 tests pass, ruff check and ruff format --check clean.

Related

Found by CodeRabbit on #1798, where the same gap is fixed for forecasting. agentic_anomaly_detection has it too — fixed on #1801, since that evaluator is not on master yet. After those three, the only ungated kind is agentic_conversation, which is ungated by design.

🤖 Generated with Claude Code

`run_agentic_what_if` already computed `pass_power_k`; the evaluator never read
it, asking `if not summary.pass_at_k` for the verdict. `--gate power` on a flaky
item therefore reported a pass on the strength of one good run out of K -- the
exact case the gate exists to catch, reporting the opposite of what happened.

The evaluator now takes `gate` like every other multi-run kind, asks
`gate_passed(...)`, and carries the gate's note into the assertion message, which
matters here because the message body describes the BEST run and under pass^K that
can be one that passed. It also stamps the gate on the run metadata and publishes
pass@K/pass^K/gate_passed, so a gated run is readable in Langfuse rather than only
in the exit code.

The default is unchanged: without --gate this is pass@K exactly as before, pinned
by a test alongside the power case. The power test fails against the previous
code.

Found by CodeRabbit on #1798, where the same gap was fixed for forecasting.
`agentic_anomaly_detection` has it too and is fixed on its own branch, #1801,
because that evaluator does not exist on master yet.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
@coderabbitai

coderabbitai Bot commented Sep 24, 2026

Copy link
Copy Markdown

Warning

Review limit reached

  • Run on-demand review

This review includes 3 billable files and costs up to $0.75.

Or wait 33 minutes for your next included review.

Check out review usage here.

View limit details

Limit details: You’ve used the included review currently available.

Learn how review limits work.

Review configuration:

⚙️ Run configuration

Configuration used: Organization UI

Review profile: CHILL

Plan: Advanced

Run ID: 95733749-672f-4bfc-996a-4e4aa392e6ff

📥 Commits

Reviewing files that changed from the base of the PR and between 9be051b and 313505a.

📒 Files selected for processing (3)
  • packages/gooddata-eval/src/gooddata_eval/cli/agentic_runner.py
  • packages/gooddata-eval/src/gooddata_eval/core/agentic/what_if.py
  • packages/gooddata-eval/tests/test_agentic_what_if.py

Comment @coderabbitai help to get the list of available commands.

Tomkess added a commit that referenced this pull request Sep 24, 2026
`run_agentic_anomaly_detection` already computed `pass_power_k`; the evaluator
never read it, asking `if not summary.pass_at_k` for the verdict, so `--gate
power` on a flaky item reported a pass on the strength of one good run out of K.

Wired the same way as the other multi-run kinds: the evaluator takes `gate`, the
dispatch passes it, `gate_passed(...)` decides, the gate's note goes into the
assertion message -- which matters because the body describes the BEST run, and
under pass^K that can be one that passed -- and pass@K/pass^K/gate_passed reach
Langfuse alongside the run metadata stamp.

The default is unchanged: without --gate this is pass@K exactly as before, pinned
by a test beside the power case. The power test fails against the previous code.

Found by CodeRabbit on #1798. Forecasting is fixed there; what-if, which is
already on master, in #1831.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
@codecov

codecov Bot commented Sep 24, 2026

Copy link
Copy Markdown

Codecov Report

✅ All modified and coverable lines are covered by tests.
✅ Project coverage is 83.08%. Comparing base (9be051b) to head (313505a).

Additional details and impacted files
@@           Coverage Diff           @@
##           master    #1831   +/-   ##
=======================================
  Coverage   83.08%   83.08%           
=======================================
  Files         330      330           
  Lines       21763    21767    +4     
=======================================
+ Hits        18082    18086    +4     
  Misses       3681     3681           

☔ View full report in Codecov by Harness.
📢 Have feedback on the report? Share it here.

🚀 New features to boost your workflow:
  • ❄️ Test Analytics: Detect flaky tests, report on failures, and find test suite problems.

Tomkess added a commit that referenced this pull request Sep 24, 2026
`run_agentic_anomaly_detection` already computed `pass_power_k`; the evaluator
never read it, asking `if not summary.pass_at_k` for the verdict, so `--gate
power` on a flaky item reported a pass on the strength of one good run out of K.

Wired the same way as the other multi-run kinds: the evaluator takes `gate`, the
dispatch passes it, `gate_passed(...)` decides, the gate's note goes into the
assertion message -- which matters because the body describes the BEST run, and
under pass^K that can be one that passed -- and pass@K/pass^K/gate_passed reach
Langfuse alongside the run metadata stamp.

The default is unchanged: without --gate this is pass@K exactly as before, pinned
by a test beside the power case. The power test fails against the previous code.

Found by CodeRabbit on #1798. Forecasting is fixed there; what-if, which is
already on master, in #1831.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
(cherry picked from commit 5632da8)
Tomkess added a commit that referenced this pull request Oct 2, 2026
…together

Rebuilt from master rather than advanced: the branch's conversation.py
predated master's multi-turn context work (QA-29448), while #1789 had
already merged it, so merging master into the old tip would have meant
hand-resolving a feature the PR branch already carried correctly.

master + #1789 #1797 #1798 #1801 #1816 #1831 #1839. Three reconciliations
the individual PRs cannot make on their own:

- Dispatch registration in cli/agentic_runner.py is additive across four
  PRs that each add an evaluator; each pair conflicts and each resolution
  is the union.
- #1816's structural test requires every multi-run evaluator to call
  build_failed_runs. dashboard_summary, forecasting and anomaly_detection
  postdate it and had no attachment point, so each grew one: a per-run
  detail function, build_failed_runs over the same predicate runs_passed
  is taken over, and failed_runs on both the outcome and the assertion
  error. dashboard_summary's _detail took the whole summary, so it is now
  a thin wrapper over a per-run _run_detail.
- #1789 adds exit_reason/turns_used while #1816 moves the same dicts
  behind _run_detail. Both land: the per-run fields go into _run_detail,
  and max_iterations stays at the item level since it is the same for
  every run.

Also supplies summary_input to #1816's failed-runs report test, which
otherwise fails a dashboard-summary item on a missing fixture field
before its evaluator is reached.

1520 passed, 1 skipped. ruff clean on everything these PRs touch; the two
pre-existing format offenders under tests/ come from master untouched.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Tomkess added a commit that referenced this pull request Oct 5, 2026
* feat(gooddata-eval): add the agentic anomaly-detection evaluator

Scores the anomaly-detection skill end to end, with the granularity map aligned to what
the service actually accepts and whole-conversation latency and cost rather than the
goal turn's alone.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* fix(gooddata-eval): let --gate decide an anomaly item, not pass@K alone

`run_agentic_anomaly_detection` already computed `pass_power_k`; the evaluator
never read it, asking `if not summary.pass_at_k` for the verdict, so `--gate
power` on a flaky item reported a pass on the strength of one good run out of K.

Wired the same way as the other multi-run kinds: the evaluator takes `gate`, the
dispatch passes it, `gate_passed(...)` decides, the gate's note goes into the
assertion message -- which matters because the body describes the BEST run, and
under pass^K that can be one that passed -- and pass@K/pass^K/gate_passed reach
Langfuse alongside the run metadata stamp.

The default is unchanged: without --gate this is pass@K exactly as before, pinned
by a test beside the power case. The power test fails against the previous code.

Found by CodeRabbit on #1798. Forecasting is fixed there; what-if, which is
already on master, in #1831.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
Tomkess added a commit that referenced this pull request Oct 5, 2026
Rebuilt on master now that #1789 (loop exit_reason) and #1801 (anomaly
detection) have landed. Carries #1797, #1798, #1816, #1831, #1839.

Three reconciliations the individual PRs cannot make on their own:

- Dispatch registration in cli/agentic_runner.py is additive across the
  evaluator PRs; each pair conflicts and each resolution is the union.
- #1789's exit_reason/turns_used are now on master in the same item-detail
  dicts #1816 moves behind a per-run builder. Both land: the per-run fields
  go into _run_detail, and max_iterations stays at the item level because it
  is the same for every run.
- #1816's structural guard requires every multi-run evaluator to call
  build_failed_runs. dashboard_summary, forecasting and anomaly_detection
  postdate it and had no attachment point, so each grew one. anomaly
  detection is now ON MASTER without it, so this gap stops being a merge
  artefact the day #1816 lands.

Also supplies summary_input to #1816's failed-runs report test, which
otherwise fails a dashboard-summary item on a missing fixture field before
its evaluator is reached.

1524 passed, 1 skipped. ruff clean.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Tomkess added a commit that referenced this pull request Oct 5, 2026
`run_agentic_dashboard_summary` already computed `pass_power_k`; the
evaluator never read it, asking `if not summary.pass_at_k` for the verdict,
and dispatch never passed `gate` through at all. So `--gate power` on a
flaky item returned a pass on the strength of one good run, and
`run_agentic_items` then recorded gate_passed=True -- the report labels the
run `power` while the item was decided by pass@K.

Same defect already fixed in forecasting (#1798), anomaly detection (#1801)
and what-if (#1831). A sweep of all twelve agentic evaluators says this was
the last one carrying it.

The gate also now reaches Langfuse: stamp_gate_metadata on the run and
log_gate_scores per trace, so the published scores agree with the verdict
instead of describing a gate the evaluator ignored.

Found in review by CodeRabbit.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Tomkess added a commit that referenced this pull request Oct 5, 2026
The same defect was found four times -- forecasting, anomaly detection,
what-if and dashboard summary -- every time by a reviewer reading the diff,
because nothing asserted the shape. An evaluator computes `pass_power_k` and
then asks `if not summary.pass_at_k` for the verdict, so `--gate power` on a
flaky item returns a pass while the report still labels the run `power`.
Nothing looks broken; a flaky item has quietly been promoted.

Structural, over every kind in AGENTIC_TEST_KINDS, skipping the ungated one.
It lives here rather than in either fixing PR because neither carries both
halves: #1797 fixes dashboard summary and #1831 fixes what-if, so the guard
only passes where both are present.

1574 passed, 2 skipped.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant