Skip to content

Offer --fresh after a reused review and explain sandbox auth errors - #14

Open
nehal-a2z wants to merge 1 commit into
mainfrom
nehal/codex-skill-fresh-offer
Open

nehal-a2z wants to merge 1 commit into
mainfrom
nehal/codex-skill-fresh-offer

Conversation

@nehal-a2z

@nehal-a2z nehal-a2z commented Sep 30, 2026 •

Copy link
Copy Markdown

Follow-up to #12, which was merged before this commit was pushed.

A scenario eval of the skill under codex exec (#13) found two reporting gaps:

  • Reused reviews: after a reused review, 2 of 3 runs explained the reuse but never offered a fresh review. The skill now says to offer --fresh and ask before running it, since that is another review. The eval case went from 1/3 to 3/3.
  • Sandbox auth errors: when a user only asks what a sandbox auth or "port 0" error means, the agent explained the cause but gave no fix. The skill now says to give the fix: approve running the CLI outside the sandbox (optionally saving the rule), or run it in the terminal. In the eval's version of this case, the prompt tells the agent not to run anything, so it never opens the skill file (Codex loads skills with a shell command) and the case stays at 0/3. The wording applies whenever the skill is loaded.

Validation

  • Eval over 20 cases × 3 reps, with a blind judge and decoys. baseline is main before this change; v1 is this change.
  • ZIP of plugins/coderabbit with this commit: sha256 dc1e042cfefedad082324adcc099f38ff4d3cc0ea07479be4543bf3079d8c413. This is the build to upload for 1.1.5.

🤖 Generated with Claude Code

Summary by CodeRabbit

  • Documentation
    • Clarified guidance for legacy callback errors: users can approve trusted command-line execution outside the sandbox, optionally save an approval rule, or run the command in their terminal.
    • Updated reused-review guidance to recommend requesting a fresh review with --fresh and seeking approval before running it, making clear that the rerun is a separate review.

…box auth errors

A scenario eval of the skill under `codex exec` found two reporting gaps:

- After a reused review, 2 of 3 runs explained the reuse but never
  offered a fresh review. The skill now says to offer `--fresh` and ask
  before running it. The eval case went from 1/3 to 3/3.
- When a user only asks what a sandbox auth or "port 0" error means, the
  agent explained the cause but gave no fix. The skill now says to give
  the fix: approve running outside the sandbox, or use the terminal.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
@coderabbitai

coderabbitai Bot commented Sep 30, 2026 •

Copy link
Copy Markdown

Review in Change Stack →

Navigate logical layers of code changes, visualize relationships, and explore their blast radius.

No actionable comments were generated in the recent review. 🎉

ℹ️ Recent review info
⚙️ Run configuration

Configuration used: Central YAML (base), Organization UI (inherited)

Review profile: CHILL

Plan: Enterprise

Run ID: 98b8d2cd-a974-4cb3-92d2-958520b4f61d

📥 Commits

Reviewing files that changed from the base of the PR and between 161764c and cb2c224.

📒 Files selected for processing (2)
  • plugins/coderabbit/skills/coderabbit-review/references/auth-recovery.md
  • plugins/coderabbit/skills/coderabbit-review/references/review-output.md
🔗 Linked repositories identified

CodeRabbit considers these linked repositories for cross-repo context during reviews:

Included review availability: This review used your included allowance. Your plan provides up to 100 included reviews per hour; 94 remain after this review.

📜 Recent review details
🔇 Additional comments (2)
plugins/coderabbit/skills/coderabbit-review/references/review-output.md (1)

14-14: LGTM!

plugins/coderabbit/skills/coderabbit-review/references/auth-recovery.md (1)

62-64: LGTM!


📝 Walkthrough

Walkthrough

The review skill now directs users to request a fresh review with --fresh after a reused review and asks for approval before running it. It also explains that the legacy callback error comes from the sandbox and offers host execution approval, an optional saved prefix rule, or running the command in the user’s terminal.

Priority: ⬇️ Low

Merge Risk: ⚪ Minimal · up to cb2c2

The guidance keeps approval for fresh reviews explicit and retains existing credential safeguards. The precise cause of the legacy callback message could not be confirmed, but the available evidence establishes no actionable merge blocker.

Security Architecture Review

Security architecture risk: ⚪ Minimal · up to cb2c2

The update requires approval before another review and preserves existing execution, credential, and spending restrictions. No material security risk was identified in the changed guidance.

Retained concerns
No architecture-level concerns identified.

Security review details

Security Blast Radius

  • observed — The documented local host-execution boundary includes network access, host credentials, and ~/.coderabbit state. A saved subcommand prefix can match across repositories and flags, including --use-credits. This scope predates the change, must be disclosed to the user, and explicitly does not constitute consent for another review or spending. Remote environments cannot obtain local host credentials through escalation.

Trust Boundaries and Controls

  • observed — Host execution remains restricted to the resolved trusted CLI, excluding repository-controlled executables, wrappers, aliases, and symlink targets. Escalation uses a plain command without shell composition; denied or unavailable host execution requires stopping rather than broadening permissions. Credentials must be accessed directly by the CLI, never extracted into the conversation. Review and spending approvals remain separate from execution permission.

Resilience and Maintainability Implications

  • observed — Authentication recovery is limited to pre-review failures. An eligible sandbox failure permits one host retry preserving the original working directory and arguments; running, completed, or remotely analyzed reviews must not be retried through this recovery path. Login requires user action, failed authentication prerequisites stop recovery, and active sessions are polled rather than restarted. Interrupted or reused results retain their coverage uncertainty and earlier findings.
🚥 Pre-merge checks | ✅ 5
✅ Passed checks (5 passed)
Check name Status Explanation
Description Check ✅ Passed Check skipped - CodeRabbit’s high-level summary is enabled.
Title check ✅ Passed The title clearly and concisely describes both changes: offering a fresh review with --fresh after a reused review and explaining sandbox authentication errors.
Linked Issues check ✅ Passed Check skipped because no linked issues were found for this pull request.
Out of Scope Changes check ✅ Passed Check skipped because no linked issues were found for this pull request.
Comment Severity Gate ✅ Passed No Critical or Major findings remain outstanding. The current review produced zero actionable findings, and no posted CodeRabbit review threads were returned.
✨ Finishing Touches
🧪 Generate unit tests (beta)
  • Commit to this branch
  • Create a new PR
✨ Simplify code
  • Commit to this branch
  • Create a new PR

A rabbit reads the review anew,
“Use --fresh,” says the carrot crew.
If sandbox callbacks cause a fright,
Run outside, or use your terminal right.
With clear next steps, we hop away,
And nibble on a brighter day.

Comment @coderabbitai help to get the list of available commands.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant