Statistical analysis for LLM evaluations, from model comparisons to inference resilient to LLM judge bias, optimized for small sample sizes. All defaults battle-tested in Monte Carlo simulations.
-
Updated
Aug 28, 2026 - Python
Statistical analysis for LLM evaluations, from model comparisons to inference resilient to LLM judge bias, optimized for small sample sizes. All defaults battle-tested in Monte Carlo simulations.
Measure prompt and skill improvements with blind A/B comparison.
Another day, another Awesome List repo. A comprehensive list of Chainforge-related content
Who evaluates the evaluator? Judicator audits LLM-as-a-Judge systems for 7 documented bias types. Zero config. Works with any LLM.
Official implementation for "GLaPE: Gold Label-agnostic Prompt Evaluation and Optimization for Large Language Models" (stay tuned & more will be updated)
The prompt engineering, prompt management, and prompt evaluation tool for Python
pi extension for fixed-task-set eval runs and prompt/system comparisons with reproducible reports
A project to take a suboptimal prompt from Langsmith, enhance it, submit it again, and then reevaluate the results. #LangSmith #PromptEngineer
The prompt engineering, prompt management, and prompt evaluation tool for TypeScript, JavaScript, and NodeJS.
Prompt Architect: reusable prompt patterns, approval-gated workflows, and verification checklists for AI-assisted development.
Production-grade prompt engineering, migration audits, and eval gates for GPT-5.6 Sol, Terra, and Luna.
A Simple Prompt Optimization Using 3 different algorithms for testing.
Audit and score Codex skills with a transparent, evidence-based custom rubric
A lightweight CLI tool for evaluating LLM prompts. Run prompts side-by-side to instantly compare token usage, tone, reading level, and API costs. Natively supports OpenAI, Anthropic, Gemini, Groq, and Ollama. Perfect for safely refactoring prompts, tuning bot personalities, or proving that "be concise" actually lowers your bill.
A few prompts that I am storing in a repo for the purpose of running controlled experiments comparing and benchmarking different LLMs for defined use-cases
Benchmark and continuously improve your Superwhisper custom modes against your own voice recording history.
Declarative LLM prompt evaluation harness, easy to use UI, all in one binary.
Local-first LLM evaluation for Ollama: benchmark, compare, judge, battle, and export results.
Test prompt variants across LLM providers with LLM-as-judge evaluation
The prompt engineering, prompt management, and prompt evaluation tool for Ruby.
Add a description, image, and links to the prompt-evaluation topic page so that developers can more easily learn about it.
To associate your repository with the prompt-evaluation topic, visit your repo's landing page and select "manage topics."