Stockholm, Sweden
I am a researcher and engineer focused on making AI systems more reliable, aligned, and capable through feedback-driven learning and structured tool use.
My work centers on Reinforcement Learning from AI Feedback (RLAIF), Model Context Protocol (MCP), test-time training (TTT), and building robust agentic systems. I am particularly interested in how models can improve during inference, how they can safely interact with external tools via standardized protocols, and how automated feedback loops can be used to align behavior without relying solely on human annotation.
I value clean interfaces, reproducible experiments, and minimal abstractions that scale.
| Focus Area | Description |
|---|---|
| RLAIF | Using AI-generated feedback to steer policy improvement and preference learning |
| MCP | Standardizing how agents and models interact with tools and external systems |
| TTT | Adapting or fine-tuning behavior at inference time for improved performance |
| Agentic Systems | Building reliable, tool-using agents that plan, execute, and recover |
- Training & evaluation: PyTorch, Jupyter
- Version control: Git
- Environment: Linux
- Writing & docs: Markdown
| Skill | Area |
|---|---|
| Reinforcement Learning | RLAIF, RLHF behavior |
| Model Context Protocol | Tool interfaces & agent integration |
| Test-Time Training | Inference-time adaptation |
| Agentic Systems | Tool use, planning, execution |
| Python | Core development & research |
- Python — primary language for research and engineering