Skip to content

Add PUMA: Semantic-Preserving Early Exit for Reasoning Models - #7

Open
ZhishanQ wants to merge 1 commit into
XiaoYee:mainfrom
ZhishanQ:add-puma
Open

Add PUMA: Semantic-Preserving Early Exit for Reasoning Models#7
ZhishanQ wants to merge 1 commit into
XiaoYee:mainfrom
ZhishanQ:add-puma

Conversation

@ZhishanQ

Copy link
Copy Markdown

Adding PUMA (arXiv 2026-05) to Efficient Reasoning during Inference → Length Budgeting at the top of the chronological list, alongside DEER ("Dynamic Early Exit in Reasoning Models").

PUMA is a plug-and-play inference-time early-exit framework for Large Reasoning Models. Unlike answer-level early-exit methods (DEER uses confidence; others use trial-answer consistency), PUMA introduces reasoning-level semantic redundancy as a complementary stopping signal — detected by a lightweight Redundancy Detector (fine-tuned Qwen3-Embedding-0.6B with a contrastive objective) — followed by an Answer Verification window confirming the candidate exit.

Results across 5 LRMs (DeepSeek-R1-Distill-Qwen-7B/14B/32B, Llama-3.1-Nemotron-Nano-8B, Qwen3-30B-A3B-Thinking) and 5 reasoning benchmarks (MATH-500, AIME24, AIME25, OlympiadBench, GPQA-Diamond): 26.2% average token reduction with preserved (slightly improved) accuracy; 1.40× / 1.28× wall-clock speedup on DS-7B/14B. Also generalizes to LiveCodeBench (code) and zero-shot VLM reasoning (MathVista, MathVision).

Thanks!

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant