Skip to content

docs: simplify benchmark results - #540

Open
dalongbao (dalongbao) wants to merge 1 commit into
microsoft:mainfrom
dalongbao:agent/average-skill-result-graphs
Open

docs: simplify benchmark results#540
dalongbao (dalongbao) wants to merge 1 commit into
microsoft:mainfrom
dalongbao:agent/average-skill-result-graphs

Conversation

@dalongbao

Copy link
Copy Markdown
Contributor

No description provided.

Copilot AI lite review requested due to automatic review settings August 14, 2026 12:24

Copilot AI left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Pull request overview

This PR updates the Skills documentation to present benchmark results in a more aggregated, easier-to-scan form by describing averaged held-out finale outcomes and removing per-budget snapshot tables, with regenerated plots to match.

Changes:

  • Update the “Performance versus overall cost” description to reflect averaging across the three held-out finale runs per treatment cell.
  • Clarify the ALFWorld baseline interpretation in the “Performance versus finale cost” section and remove the detailed $5/$10/$25 snapshot tables.
  • Regenerate the three “overall cost” SVG plots to reflect the simplified/aggregated presentation.

Reviewed changes

Copilot reviewed 1 out of 4 changed files in this pull request and generated no comments.

File Description
skills/README.md Updates benchmark narrative to describe averaged points and removes detailed budget snapshot tables.
skills/assets/agent-lightning-spreadsheetbench-accuracy-overall-cost.svg Regenerated overall-cost plot consistent with the updated aggregation description.
skills/assets/agent-lightning-officeqa-correctness-overall-cost.svg Regenerated overall-cost plot consistent with the updated aggregation description.
skills/assets/agent-lightning-alfworld-success-overall-cost.svg Regenerated overall-cost plot consistent with the updated aggregation description.

💡 Add a code-review agent skill or configure MCP servers for context-aware, tailored reviews. Learn more in the docs.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants