Kangrui Wang*, Pingyue Zhang*, Zihan Wang*, Yaning Gao*, Linjie Li*, Qineng Wang, Hanyang Chen, Chi Wan, Yiping Lu, Zhengyuan Yang, Lijuan Wang, Ranjay Krishna, Jiajun Wu, Li Fei-Fei, Yejin Choi, Manling Li
(* equal contribution)
VAGEN is a reinforcement learning (RL) framework that trains multi-turn VLM agents (vision-language model agents) to build an internal world model through explicit visual state reasoning. Instead of rewarding only task success, VAGEN reinforces the agent's world model reasoning itself, decomposed into StateEstimation ("what is the current state?") and TransitionModeling ("what comes next?"), with a turn-level WorldModeling Reward (LLM-as-Judge) and Bi-Level GAE for turn-aware credit assignment. Combining world models with reinforcement learning, a 3B VLM trained with VAGEN scores 0.82 across five visual agent benchmarks, a 3x improvement over its untrained backbone (0.21), outperforming GPT-5 (0.75), Gemini 2.5 Pro (0.67), and Claude 4.5 (0.62).
We introduce VAGEN, a multi-turn reinforcement learning framework designed specifically for training vision-language model (VLM) agents. Built upon this framework, we propose World Modeling RL, a novel reinforcement learning approach that significantly improves the multi-turn performance of VLMs by explicitly supervising their worldmodel reasoning process, as shown in Figure 1.
We frame multi-turn VLM agentic tasks as a Partially Observable Markov Decision Process (POMDP), shown in Figure 2.
![]() |
![]() |
|---|---|
| Figure 1. Overview of the VAGEN framework. | Figure 2. POMDP formulation of multi-turn VLM agentic tasks. |
[2026/08]: Added support for Verl 0.9.0, introduced compaction RL, and decoupled the harness layer.
- Compaction RL — a multi-turn paradigm alongside concat and no-concat. Turns accumulate until a token budget is reached, then the conversation is summarised and reopened from that summary, so an episode longer than the context window still trains as one trajectory. See Multi-turn Compacted Training. Reference: (CompactionRL)
- Environment, harness, training backend and algorithm are decoupled. Anything subclassing
BaseEnvorBaseHarnessplugs into both training and evaluation without either being modified — see Custom Harness and Custom Environment.
[2026/02] We have migrated the main branch to VAGEN-Lite, a lightweight and clean reimplementation built on VERL agent-loop for easy customization and stable performance. For the previous full-featured release, please visit the vagen-legacy branch.
[2025/12] Introducing VAGEN-Lite: a lightweight and clean reimplementation of VAGEN, built on the VERL agent-loop for easy customization and stable performance.
[2025/09] VAGEN is accepted by Neurips 2025
[2025/04] We've introduced a new modular design for environments and services in VAGEN:
- Enhanced environment framework for easier creation of custom environments
- New service architecture for efficient distributed training
- Check out our new guides:
- Creating Environments: New environment protocol.
- Creating Services: We now support hosting environments in a separate process
[2025/03] We release VAGEN, a multi-turn reinforcement learning framework for training VLM Agents!
conda create -n vagen python=3.12 -y
conda activate vagen
git clone --recursive --branch release-ready https://github.com/JamesKrW/VAGEN.git
cd VAGEN
bash scripts/install.shscripts/install.sh fetches the pinned verl submodule, installs VAGEN with a rollout
engine, then verl, and checks the result. It is idempotent, so it is safe to re-run.
SKIP_ENGINE=1 installs VAGEN without an engine if you already have one.
vLLM is the default and the verified training path. For SGLang evaluation and local serving:
BACKEND=sglang bash scripts/install.shInstalling the SGLang extra does not switch the shipped training launchers: they source
baseline_vllm.flags and still select vLLM. Use the SGLang evaluation launchers where one
is provided; a training launcher needs explicit SGLang rollout configuration and its own
model-level validation.
Use one engine per environment. They are mutually exclusive, and not by preference:
each pins a different flashinfer patch version, so pip refuses to install them together.
Use two conda environments if you want both.
Doing it by hand
git submodule update --init --recursive # verl, pinned; the scripts will not run without it
pip install -e ".[vllm]" # or ".[sglang]" -- pick one, never both
pip install --no-deps -e ./verl # --no-deps: verl's pins would undo the line above
pip install accelerate codetiming datasets dill hydra-core numpy pandas peft pyarrow \
pybind11 pylatexenc ray tensordict torchdata wandbThe engine, torch and transformers versions all live in setup.py's
extras_require, so there is one place that says which versions go together.
No flash-attn step: it publishes no wheel past torch 2.9, so on a newer torch installing
it means a source build. transformers[kernels], which the extras pull in, instead fetches
a prebuilt kernels-community/flash-attn2 from the Hub on first use.
verl is imported from the checkout rather than from PyPI, and the training scripts find
it at VAGEN/verl (the submodule) or ../verl (a sibling checkout), in that order. Set
VERL=/path/to/verl to override.
Environments in this repository — vagen/configs/env_registry.yaml is the list that
matters: Sokoban, FrozenLake, SpatialGym, PrimitiveSkill (ManiSkill), and
RemoteEnv, which is how Navigation runs. The five benchmarks pictured above are the
paper's; SVG is not part of this release.
Some need their own setup: spatial_gym (dataset
download, plus matplotlib and scipy from its requirements.txt — without them the
registry drops the environment and you get KeyError: Unknown env name: SpatialGym),
navigation (AI2-THOR),
primitive_skill (ManiSkill).
wandb login
# Self-hosted W&B:
# WANDB_BASE_URL=https://your-wandb-host wandb login --host https://your-wandb-hostcd VAGEN
# Qwen2.5-VL: concat / no-concat / compact
bash examples/train/sokoban/train_default_gae_qwen25vl3b.sh
bash examples/train/sokoban/train_ppo_no_concat_qwen25vl3b.sh
bash examples/train/sokoban/train_default_gae_compact_qwen25vl3b.sh
# Qwen2.5-VL: bi-level GAE with state reward
bash examples/train/sokoban/train_bi_level_gae_sr_qwen25vl3b.sh
# Qwen3-VL and Qwen3.5
bash examples/train/sokoban/train_default_gae_qwen3vl4b.sh
bash examples/train/sokoban/train_default_gae_qwen35_4b.sh
# InternVL3.5 and GLM-4.6V-Flash; HARNESS=concat|no_concat|compact
HARNESS=concat bash examples/train/sokoban/train_default_gae_internvl35_2b.sh
HARNESS=concat bash examples/train/sokoban/train_default_gae_glm46v_flash.sh
# Validation only, without starting training
bash examples/train/sokoban/train_default_gae_internvl35_2b.sh \
trainer.val_only=true trainer.save_freq=-1 trainer.test_freq=-1See Configuration for harnesses, estimators, model-specific flags, and state-reward settings.
cd VAGEN
# Local vLLM
MODEL_PATH=Qwen/Qwen2.5-VL-3B-Instruct \
bash examples/evaluate/sokoban/vllm/eval_qwen25_vl_3b.sh
# OpenAI-compatible endpoint
bash examples/evaluate/sokoban/run_eval.shSee Evaluation for other environments and backends.
With sglang instead
Requires the sglang extra, which is mutually exclusive with vLLM — see Installation.
bash examples/evaluate/frozenlake/sglang/eval_qwen25_vl_3b.shTo train on your own environment, follow the steps below.
-
Use
GymImageEnvas the base class: -
Refer to Sokoban for a full implementation example:
Add your environment entry to:
vagen/configs/env_registry.yamlPrepare training and validation configs:
train.yamlval.yaml
You can follow the Sokoban examples as templates:
Write your training script based on:
Add an estimator under vagen/custom_advantage/ and import its module from
vagen/custom_advantage/__init__.py:
from vagen.custom_advantage import AdvantageInputs, AdvantageOutputs, advantage_estimator
@advantage_estimator("my_estimator", needs_critic=True)
def my_estimator(inputs: AdvantageInputs):
returns = inputs.rewards
advantages = (returns - inputs.values) * inputs.response_mask
return AdvantageOutputs(advantages=advantages, returns=returns)Select it in a training command with algorithm.adv_estimator=my_estimator. See
inputs.py for the available inputs and
trajectory_algos.py for complete examples.
A harness decides what the model sees on each turn: whether the next call continues the
conversation so far, or starts a fresh one, and what that fresh one begins with. concat,
no_concat and compact are three different answers to that.
Anything that subclasses BaseEnv or BaseHarness works in both training and evaluation
without changing either. That is possible because a harness is deliberately small: it has
no tokenizer, no client and no environment, and it never touches tokens, masks or rewards.
All it does is decide what goes into the next call. So the same harness can drive a training
rollout and an evaluation run against a closed API.
A harness also doesn't assume the conversation is kept on the client side. That leaves room
to add a backend that keeps it on the server instead (OpenAI's previous_response_id, an
SGLang session, a vLLM prefix cache) without rewriting any harness. No backend shipped today
does that — they all re-send the whole message list every turn.
from vagen.core.harness import BaseHarness, Call
from vagen.harness import register_harness
@register_harness("mine")
class MyHarness(BaseHarness):
#: Whether one episode can end up in more than one row. The trainer asks the harness
#: rather than keeping a list of the ones it knows, and pairs the estimator accordingly.
splits_episode_across_rows = True
def next_call(self) -> Call: ... # the only required methodTraining and evaluation both read a harness key, and both accept either a registered name
or an import path. What differs is who imports your module:
# training -- vagen/configs/vagen_multiturn.yaml, or a -o override
trainer:
harness: mine
# verl builds a registry per worker process, so the module has to be imported inside each
# one. That is what this is for; without it the decorator never runs and the name is
# unknown:
# actor_rollout_ref.model.external_lib=mypkg.harnesses# evaluation -- examples/evaluate/<env>/config.yaml
envs:
- name: Sokoban
harness: mypkg.harnesses:MyHarness # an import path, not a bare nameIn evaluation, use the import path. The @register_harness decorator only runs if
something imports your module, and run_eval has no external_lib setting to make that
happen — so a bare harness: mine fails with:
unknown harness 'mine'; choose from ['compact', 'concat', 'no_concat']
Give the import path instead and VAGEN imports the module itself. This is why the import path is supported at all: a new harness is usually tried in evaluation first, where everything is configured in yaml and no decorator has had a chance to run.
Full contract and the budget hooks: vagen/core/harness.py; the
three implementations: vagen/harness/.
See the Documentation for more customization options:
- Custom Filter — Trajectory filtering (e.g., Reward Variance (RV) filter in RAGEN)
- Custom Metric - Add W&B logging metrics
- Configuration - Training configuration reference
refer to vagen/configs/vagen_multiturn.yaml
# Enable no concat mode: input is system prompt + current step observation
trainer:
harness: no_concat # concat | no_concat | compact
# no_concat and compact put one episode in several rows, so the advantage estimator has
# to be one that stitches them back together. verl's own `gae`/`grpo` score a row at a
# time and would drop every turn's credit at the row boundary; the trainer refuses that
# pairing at startup rather than training on it.
algorithm:
adv_estimator: default_gae # or bi_level_gae | turn_level_gae | token_level_gae
# | trajectory_grpo
# default_gae is the vanilla baseline: the episode's whole reward lumped onto its
# last token, which is what single-turn RLHF does. It stitches rows like the others,
# so it stays comparable under no_concat and compact where verl's `gae` would not.
# Warning:
# - If you set a training-data rollout dir AND enable image logging, training images will also be dumped to disk.
# This can consume a large amount of storage very quickly. Monitor disk usage and consider cleanup/limits.
trainer:
log_image:
enable: false # true can enable saving rollout/validation images to disk
max_pending: 2 # max concurrent async image dump tasks
png_compress_level: 0 # PNG compression (0 = fastest, 9 = smallest)# export HF_TOKEN=xxx
huggingface_hub:
hf_save_freq: null # upload every N steps (must be a multiple of trainer.save_freq); null = disabled
repo_id: vagen-training # the shipped default; enabling upload with it unchanged
# pushes to a repo of that name under your account
private: false filter:
name: reward_variance_top_p # refer to vagen/custom_filter
filter_kwargs:
top_p: 0.9
enable: False # set to true to enable filtering, recommended for grpo traininingSee docs/issues.md
If you find our framework and paper useful, we appreciate it if you could cite our work:
@inproceedings{wang2025vagen,
title={VAGEN: Reinforcing World Model Reasoning for Multi-Turn VLM Agents},
author={Kangrui Wang and Pingyue Zhang and Zihan Wang and Yaning Gao and Linjie Li and Qineng Wang and Hanyang Chen and Chi Wan and Yiping Lu and Zhengyuan Yang and Lijuan Wang and Ranjay Krishna and Jiajun Wu and Li Fei-Fei and Yejin Choi and Manling Li},
booktitle={The Thirty-ninth Annual Conference on Neural Information Processing Systems},
year={2025},
url={https://arxiv.org/abs/2510.16907}
}










