Model Details
Table with columns: Item, Description| Item | Description |
|---|
| Base model | Qwen/Qwen2.5-1.5B-Instruct |
| Architecture | Qwen2ForCausalLM |
| Checkpoint format | BF16 Safetensors |
| Configured context length | 32,768 tokens |
| Target environment | ALFWorld / ALFRED text environment |
| Post-training | Agent-G2 with GRPO |
| Language | English |
| Required output format | <think>...</think><action>...</action> |
Although the tokenizer metadata contains a larger generic maximum length, the model
configuration declares 32,768 positions and the released training recipe uses at most
4,096 prompt tokens plus 512 response tokens.
Evaluation
The Agent-G2 project reports the following ALFWorld success rates for this 1.5B
checkpoint:
Table with columns: Task group, Success rate| Task group | Success rate |
|---|
| Pick | 96.8% |
| Look | 100.0% |
| Clean | 100.0% |
| Heat | 92.9% |
| Cool | 84.2% |
| Pick Two | 94.7% |
| All tasks | 95.3% |
Expert-prefix guidance is enabled during training but disabled during validation in
the released configuration (gmsv.apply_on_validation=false). The reported results
therefore do not require an expert trajectory at inference time.
These results are reported by the
Agent-G2 repository and have not been
independently reproduced in this model card. Evaluation variance is not currently
available. Results may vary with the ALFWorld version, task split, prompt template,
action history, random seed, and decoding configuration.
Intended Use
This checkpoint is intended for:
- reproducing Agent-G2 results in the ALFWorld text environment;
- research on long-horizon language agents and agentic reinforcement learning;
- studying adaptive expert-prefix guidance;
- evaluating action selection over an environment-provided admissible action set.
For faithful evaluation, use the ALFWorld environment, prompt template, action parser,
and rollout loop provided by the
Agent-G2 repository. A standalone generation
only demonstrates that the checkpoint loads successfully; it does not reproduce the
interactive benchmark.
Environment Interface
At every environment step, provide the task, current observation, recent history, and
admissible actions. The released parser expects English output containing reasoning
and one action selected from the current admissible set:
<think>Reason about the observation and admissible actions.</think>
<action>put apple 1 in/on fridge 1</action>
Missing tags or outputs containing Chinese characters are marked invalid by the
released ALFWorld parser. The action inside <action>...</action> must match an action
that the environment currently allows.
Quick Start
pip install -U transformers accelerate torch
The following example performs one ALFWorld-style generation step. Replace the
placeholders with state supplied by the environment:
import torch
from transformers import AutoModelForCausalLM, AutoTokenizer
model_id = "xiamoent/Agent-G2-alfworld-1.5b"
tokenizer = AutoTokenizer.from_pretrained(model_id)
model = AutoModelForCausalLM.from_pretrained(
model_id,
torch_dtype="auto",
device_map="auto",
)
model.eval()
task_description = "<ALFWorld task>"
current_observation = "<current observation>"
admissible_actions = ["<admissible action 1>", "<admissible action 2>"]
actions_text = ", ".join(admissible_actions)
prompt = f"""
You are an expert agent operating in the ALFRED Embodied Environment.
Your task is to: {task_description}
Your current observation is: {current_observation}
Your admissible actions of the current situation are: [{actions_text}].
Now take one action. Enclose your reasoning within <think> </think> tags, then
present one admissible action within <action> </action> tags.
""".strip()
inputs = tokenizer.apply_chat_template(
[{"role": "user", "content": prompt}],
tokenize=True,
add_generation_prompt=True,
return_dict=True,
return_tensors="pt",
).to(model.device)
with torch.inference_mode():
output_ids = model.generate(
**inputs,
max_new_tokens=512,
do_sample=True,
temperature=0.4,
top_p=0.8,
top_k=20,
repetition_penalty=1.1,
)
new_tokens = output_ids[0, inputs["input_ids"].shape[-1]:]
response = tokenizer.decode(new_tokens, skip_special_tokens=True)
print(response)
The released checkpoint's generation_config.json defaults to temperature 0.7.
The example uses temperature 0.4 to match the released validation configuration.
Training
Agent-G2 uses expert ALFWorld trajectories as prefix guidance during training,
followed by policy rollouts and GRPO updates. The guidance depth is sampled per task
from a Gaussian distribution estimated online from existing rollout statistics.
Prefix guidance is a training mechanism; it is not required for validation or
deployment.
The expert-prefix store contains 3,553 ALFWorld trajectories with action lengths from
3 to 15. The released training recipe specifies:
Table with columns: Configuration, Value| Configuration | Value |
|---|
| Learning rate | 1e-6 |
| Training batch size | 16 |
| Rollouts per task | 8 |
| Maximum prompt length | 4096 |
| Maximum response length | 512 |
| Maximum ALFWorld steps | 20 |
See the paper-locked
run_alfworld.sh
for the complete recipe. The public repository does not identify the exact checkpoint
step or selection rule used for this Hub upload, so the table documents the released
recipe rather than claiming that this artifact is the final epoch checkpoint.
Limitations
- The model is specialized for the text-based ALFWorld environment and may not
generalize to other simulators or physical environments.
- It can produce malformed or inadmissible actions; environment-side validation is
required.
- Performance is sensitive to prompt formatting, observation history, decoding
settings, random seed, and environment configuration.
- The reported evaluation does not include variance across repeated runs.
- The model may inherit factual errors, biases, and other limitations from the base
model and training data.
- This checkpoint should not directly control physical systems or be used for
consequential real-world actions without independent safety mechanisms.
Citation
If you find this checkpoint useful, please cite Agent-G2:
@misc{wang2026agentg2gaussianguidanceagentic,
title={Agent-G$^2$: Gaussian Guidance for Agentic Reinforcement Learning},
author={Zixuan Wang and Yanrui Miao and Zhengxi Lu and Teng Pan and Yiwen Qiu and Hongxing Li and Peng Qiu and Ruiqing Zhang and Yongliang Shen},
year={2026},
eprint={2608.23318},
archivePrefix={arXiv},
primaryClass={cs.AI},
url={https://arxiv.org/abs/2608.23318},
}
The paper has been accepted to the EMNLP 2026 Main Conference. A public paper link
will be added when available.
Acknowledgements
Agent-G2 builds on verl-agent,
veRL, and
ALFWorld.