[ICLR 2025] General-purpose activation steering library
-
Updated
Sep 18, 2025 - Python
[ICLR 2025] General-purpose activation steering library
KV Cache Steering for Inducing Reasoning in Small Language Models
OpenAI-compatible server with a live view into any HF model's residual stream. pip install brainscope
Agents that talk through model internals — activations & J-space — instead of text. No words pass between them. The lab on top of brainscope + hidden-directions.
[ACL 2026] - Official repo for the paper: "Selective Steering: Norm-Preserving Control Through Discriminative Layer Selection"
Activation steering and trait monitoring for HuggingFace transformers
[Under Review] Not All Tokens Are Equally Useful for Steering: Robust Directions and Prefix Steering
A bilingual awesome list for refusal suppression research: benchmarks, papers, tools, models, and ecosystem updates.
[ICLR 2026] ASGuard: Activation-Scaling Guard to Mitigate Targeted Jailbreaking Attack
Official code for "Activation Steering for Accent Adaptation in Speech Foundation Models" (Interspeech 2026). Parameter-free accent adaptation via mean-shift steering vectors — no weight updates, consistent WER reductions across 8 accents.
GEMS: Geometric Constraints Enable Multi-Semantic Superposition in LLMs
Does concept-injection introspection emerge with scale? A faithful, controlled reproduction charted across model-size ladders.
🏆[ICML 2026 Spotlight] Official implementation of "DecodeShare: Tracing the Shared Subspace of LLM Decode-Time Decisions"
The paper list related to activation steering
Steer2Adapt: data-efficient inference-time LLM adaptation by composing steering vectors via Bayesian optimization over a semantic prior subspace.
Steering vectors with receipts: make one, catch one, deploy a calibrated one. pip install hidden-directions
CUDA-graph-safe per-request activation steering plugin for vLLM. pip install hotwire-vllm
Phase-aware LLM activation steering and linear probing. A memory-efficient, practical implementation of Representation Engineering (RepE) for safety research.
Reproduce Emotion Concepts and their Function in a Large Language Model on Qwen 3.6 27b.
Accepted at 19th Conference of the European Chapter of the Association for Computational Linguistics, 2026
Add a description, image, and links to the activation-steering topic page so that developers can more easily learn about it.
To associate your repository with the activation-steering topic, visit your repo's landing page and select "manage topics."