AuditPilot: auditable enterprise AI agents for evidence-grounded workflows, governed tools, evaluation harnesses, human review, and remediation delivery.
-
Updated
Jul 28, 2026 - Python
AuditPilot: auditable enterprise AI agents for evidence-grounded workflows, governed tools, evaluation harnesses, human review, and remediation delivery.
The open-source MultiAgentOps evaluation and verification harness for any industry business workflow.
An end-to-end framework for running, sandboxing, and scoring agentic LLMs on complex data-science and econometric replication tasks.
Detecting Relational Boundary Erosion in AI systems. A framework for testing whether models maintain honest, calibrated, and appropriate boundaries.
VLA ≠ VLM. Side-by-side viewer running NVIDIA Alpamayo R1 (vision-language-action) alongside Qwen2.5-VL (vision-language) on the same 44-sec SF dashcam clip at 5 Hz. 220 paired traces. Surfaces what an action-trained model sees that a scene-trained model doesn't, and vice versa.
An LLM agent over crash telemetry that must cite the tool result behind every number it states - with an automatic citation checker and eval harness.
RAG service that treats abstention as a feature: cited answers, a faithfulness gate, and a measured coverage-vs-false-answer curve. Chunking x retriever evaluation grid vs planted ground truth shows why retrieval metrics alone mislead. From-scratch BM25, LSA + RRF hybrid, FastAPI, MLflow, 31 tests, fully offline CI.
Independent, cross-tool, register-split benchmark of anti-AI-slop rulesets. It measures rules, not writers.
Wellness verification harness for companion AI. Multi-turn adversarial suites grounded in six decades of mental-health research and current clinical standards (988, VERA-MH) and law (SB 243). Point it at any chat endpoint — get an evidence-backed, reproducible report.
Autonomous financial research agent combining live market data, financial news, sentiment analysis, and private RAG with transparent execution.
Does a CLAUDE.md actually change how Claude behaves? An ablation harness: run adversarial traps with the rules and without them, grade blind, and test whether the difference is real.
Towards Evaluation Engineering: An Empirical Study of ML Evaluation Harnesses in the Wild
Constitutional governance platform for multi-agent AI systems — 261 personas, 17 divisions, a judiciary, RBAC, and a written constitution
AI content engine using an anxiety-indexed behavioral science KB, multi-stage LangGraph pipeline, and calibrated LLM-as-judge evaluation harness
Enterprise RAG lab using AWS Bedrock, Snowflake, MuleSoft, Python, and an evaluation harness for regulated lending scenarios.
An LLM-powered training-evaluation platform that scores open-ended scenario responses 0 to 10 against rubrics, with an evaluation harness that benchmarks the AI scorer against human-labelled scores.
Closed-loop LLM factory in one monorepo: data pipeline, trainer, eval-gated checkpoint promotion, serving, and an agent whose single tool is a self-extending CLI.
Turnkey — Bring Your Own Detector. A development harness for jailbreak detection.
Chaos engineering for multi-agent LLM meshes — which topology survives which failure? Runs on Nebius Serverless: CPU Job sweeps against a vLLM Endpoint, per-trial Object Storage checkpoints, kill-and-recover by design.
DoE Project
Add a description, image, and links to the evaluation-harness topic page so that developers can more easily learn about it.
To associate your repository with the evaluation-harness topic, visit your repo's landing page and select "manage topics."