File size: 11,961 Bytes
1b0fb43
 
d8883ad
 
 
 
 
 
 
 
1b0fb43
d8883ad
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
92fa0b6
d8883ad
 
 
 
 
 
 
 
 
 
 
 
 
 
92fa0b6
d8883ad
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
d9c0075
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
d8883ad
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
35
36
37
38
39
40
41
42
43
44
45
46
47
48
49
50
51
52
53
54
55
56
57
58
59
60
61
62
63
64
65
66
67
68
69
70
71
72
73
74
75
76
77
78
79
80
81
82
83
84
85
86
87
88
89
90
91
92
93
94
95
96
97
98
99
100
101
102
103
104
105
106
107
108
109
110
111
112
113
114
115
116
117
118
119
120
121
122
123
124
125
126
127
128
129
130
131
132
133
134
135
136
137
138
139
140
141
142
143
144
145
146
147
148
149
150
151
152
153
154
155
156
157
158
159
160
161
162
163
164
165
166
167
168
169
170
171
172
173
174
175
176
177
178
179
180
181
182
183
184
185
186
187
188
189
190
191
192
193
194
195
196
197
198
199
200
201
202
203
204
205
206
207
208
209
210
211
212
213
214
215
216
217
218
219
220
221
222
223
224
225
226
227
228
229
230
231
232
233
234
235
236
237
238
239
240
241
---
license: apache-2.0
pipeline_tag: text-generation
tags:
  - agents
  - tool-use
  - coding
  - reasoning
  - runtime-learning
  - on-policy-distillation
---

# O2-9B-Preview

> A 9B agentic model trained from interactions generated inside runtimes deployed across real office and software-development workflows.

## Preliminary Evaluation

O2-9B-Preview was evaluated across eleven reasoning, tool-use, and coding benchmarks. We report its same-size Qwen3.5-9B base alongside larger O2 checkpoints and cross-scale reference models. Scores are shown on their original scales, and higher is better. The average is the unweighted mean of the eleven displayed benchmarks.

<table>
  <thead>
    <tr>
      <th></th>
      <th></th>
      <th colspan="3" align="center">Reasoning</th>
      <th colspan="6" align="center">Tool Use</th>
      <th colspan="2" align="center">Coding</th>
    </tr>
    <tr>
      <th align="left">Model</th>
      <th align="center">Avg</th>
      <th align="center">GPQA</th>
      <th align="center">HMMT<br>2025</th>
      <th align="center">HMMT<br>2026</th>
      <th align="center">bfcl v3</th>
      <th align="center">bfcl v4</th>
      <th align="center">Tau2<br>Retail</th>
      <th align="center">Tau2<br>Airline</th>
      <th align="center">Tau2<br>Telecom</th>
      <th align="center">Tau3<br>Banking</th>
      <th align="center">Agent<br>Bench OS</th>
      <th align="center">SWE-<br>Verified</th>
    </tr>
  </thead>
  <tbody>
    <tr>
      <td>DeepSeek<br>V4 Flash<br>(284b–a13b)</td>
      <td align="center">71.84</td>
      <td align="center"><strong>90.40</strong></td>
      <td align="center"><strong>93.33</strong></td>
      <td align="center">87.88</td>
      <td align="center">48.50</td>
      <td align="center">57.60</td>
      <td align="center">85.09</td>
      <td align="center"><strong>90.00</strong></td>
      <td align="center"><u>95.61</u></td>
      <td align="center"><strong>22.68</strong></td>
      <td align="center">50.00</td>
      <td align="center"><strong>69.20</strong></td>
    </tr>
    <tr>
      <td>MiniMax-<br>M2.5</td>
      <td align="center">62.33</td>
      <td align="center">85.9</td>
      <td align="center">75.60</td>
      <td align="center">73.70</td>
      <td align="center">37.00</td>
      <td align="center">49.27</td>
      <td align="center">79.24</td>
      <td align="center">70.67</td>
      <td align="center">90.94</td>
      <td align="center">8.93</td>
      <td align="center">41.00</td>
      <td align="center">73.40</td>
    </tr>
    <tr>
      <td>Qwen<br>3.5–397B–<br>A17B</td>
      <td align="center">70.56</td>
      <td align="center">86.9</td>
      <td align="center">90.00</td>
      <td align="center">81.80</td>
      <td align="center">58.90</td>
      <td align="center">63.66</td>
      <td align="center">87.13</td>
      <td align="center">85.33</td>
      <td align="center">93.57</td>
      <td align="center">12.37</td>
      <td align="center">51.30</td>
      <td align="center">65.20</td>
    </tr>
    <tr>
      <td>Qwen<br>3.5–<br>27B</td>
      <td align="center">69.97</td>
      <td align="center">86.04</td>
      <td align="center">90.00</td>
      <td align="center">81.80</td>
      <td align="center">58.10</td>
      <td align="center">62.15</td>
      <td align="center">85.96</td>
      <td align="center">82.67</td>
      <td align="center">90.94</td>
      <td align="center">15.46</td>
      <td align="center">56.90</td>
      <td align="center">59.60</td>
    </tr>
    <tr>
      <td>Qwen<br>3.5–<br>9B</td>
      <td align="center">60.27</td>
      <td align="center">80.81</td>
      <td align="center">77.78</td>
      <td align="center">67.17</td>
      <td align="center">48.70</td>
      <td align="center">62.36</td>
      <td align="center">83.33</td>
      <td align="center">78.67</td>
      <td align="center">95.32</td>
      <td align="center">5.50</td>
      <td align="center">40.97</td>
      <td align="center">22.40</td>
    </tr>
    <tr>
      <td colspan="13">&nbsp;</td>
    </tr>
    <tr>
      <td>O2-9B-Preview<br>(Ours)</td>
      <td align="center">64.07</td>
      <td align="center">87.37</td>
      <td align="center">84.44</td>
      <td align="center">69.70</td>
      <td align="center">61.90</td>
      <td align="center">51.83</td>
      <td align="center">86.26</td>
      <td align="center">82.67</td>
      <td align="center"><strong>96.20</strong></td>
      <td align="center">9.28</td>
      <td align="center">45.14</td>
      <td align="center">30.00</td>
    </tr>
    <tr>
      <td>O2-27B-Preview<br>(Ours)</td>
      <td align="center"><strong>73.40</strong></td>
      <td align="center"><strong>90.40</strong></td>
      <td align="center"><strong>93.33</strong></td>
      <td align="center"><strong>90.01</strong></td>
      <td align="center"><strong>63.38</strong></td>
      <td align="center"><strong>67.10</strong></td>
      <td align="center"><strong>89.47</strong></td>
      <td align="center"><u>86.00</u></td>
      <td align="center"><u>95.61</u></td>
      <td align="center"><u>16.48</u></td>
      <td align="center"><strong>63.19</strong></td>
      <td align="center">52.40</td>
    </tr>
  </tbody>
</table>

O2-9B-Preview improves ten of the eleven reported point estimates over the base model. In a separate comparison of complete post-training pipelines built on the same 9B base, it achieves the highest reported aggregate score: 64.07, compared with 61.08 for PowerOPD, 60.22 for GRPO, 57.05 for SFT, and 54.59 for CurryOPD.

These results are preliminary. They come from an author-provided evaluation report without repeated seeds, uncertainty intervals, or matched training and runtime-compute budgets. They characterize the released checkpoint but do not isolate the causal contribution of any individual training component.

## Overview

O2-9B-Preview is a 9B-parameter preview model developed by Ant International and initialized from Qwen3.5-9B. It is designed for agentic reasoning, tool use, and coding.

The defining feature of O2-9B-Preview is where its post-training experience comes from. Instead of relying only on static prompt-response datasets, we deployed an executable agent runtime in real operational environments—including international merchant customer-service workflows and the daily development workflows of dozens of engineers—and trained the model from the resulting interactions.

Within these runtimes, the model acts on a task, observes the consequence, receives an executable verification result, and, when necessary, repairs its action. These student-visited states are preserved as structured training records and distilled back into the model. O2-9B-Preview is therefore trained not only on what a successful answer looks like, but also on what happened when an action was executed, why it failed, and how it was repaired.

## What Makes O2 Different

### Trained where agents actually work

The training loop is connected to deployed office and development environments rather than an isolated text-only data pipeline. It captures interactions with tools, repositories, tests, workflow state, and domain rules under the same kinds of conditions in which agents are expected to operate.

### Learning from executable outcomes

For each task, the runtime constructs a task-conditioned verifier suite that checks observable evidence such as tool results, environment state, code execution, repository tests, schema constraints, and domain rules. A verifier pass is treated as a certificate under a specific, versioned suite—not as unrestricted ground truth.

### Preserving failures and repairs

The runtime records the task context, model action, executed consequence, verifier outcome, localized failure feedback, and repaired action when one exists. Accepted, rejected, repaired, and unresolved attempts remain available for learning instead of discarding everything except the final answer.

### Runtime on-policy distillation

O2-9B-Preview is post-trained with Runtime OPSD, an on-policy sequence-distillation recipe over states visited by the student in the deployed runtime. Oracle outcomes anchor complete trajectories, while teacher and student score tokens under the same task context, runtime state, and generated prefix. This keeps the learning signal aligned with the information available to the model at inference time.

## Training Loop

1. A task and its runtime profile are compiled into a versioned verifier suite.
2. The model acts in the executable task environment.
3. The runtime checks the resulting state and returns an auditable outcome.
4. Failed obligations produce localized, failure-only feedback for bounded repair.
5. The interaction is stored as a versioned training record, including failures and repair transitions.
6. Runtime OPSD distills these student-visited records into the next model checkpoint.

This process turns deployed task environments into a renewable source of grounded training experience.

## Model Details

| Field | Description |
| --- | --- |
| Model | O2-9B-Preview |
| Developer | Ant International |
| Parameters | 9B |
| Base model | Qwen3.5-9B |
| Training paradigm | Verifier-guided runtime data synthesis and Runtime OPSD |
| Primary capability groups | Reasoning, tool use, and coding |
| Release stage | Preview |

## Intended Use

O2-9B-Preview is intended for research and evaluation involving:

- agentic reasoning over multi-step tasks;
- tool-using assistants operating in controlled environments;
- software-development and repository workflows;
- office and business-process automation with explicit validation; and
- research on runtime learning, executable verification, and on-policy distillation.

For state-changing or business-critical actions, the model should be deployed with sandboxed tools, least-privilege access, independent checks, and human review.

## Limitations

- **Preview evidence.** The current results are single-report point estimates without multi-seed uncertainty or matched-compute controls.
- **Capability-specific regression.** Despite a higher aggregate score, O2-9B-Preview scores lower than its base model on BFCL V4 in the reported evaluation.
- **Partial verification.** Executable verifiers encode partial specifications. They can miss valid behavior, accept incorrect behavior, or reward shortcuts when checks are incomplete.
- **Runtime dependence.** Performance can depend on tool availability, verifier coverage, environment state, and the compatibility of the deployment runtime with the training setup.
- **Operational claims.** Deployment of the training runtime in merchant-service and engineering workflows establishes real-world operational scope, not a controlled improvement in business outcomes or developer productivity.
- **Reproducibility.** Exact data mixtures, teacher checkpoint, optimizer schedule, final Runtime OPSD hyperparameters, generated-token budget, runtime-call budget, decoding settings, and immutable evaluator identifiers are not included in this preview.

## Safety and Responsible Deployment

O2-9B-Preview can generate incorrect content, malformed tool calls, insecure code, or actions that satisfy an incomplete verifier without satisfying the user's broader intent. Do not treat a verifier pass as proof of unrestricted correctness. Deployers should use held-out checks, audit false acceptances and rejections, isolate executable environments, minimize access to sensitive data, and require human confirmation for consequential actions.

The model has not been established as suitable for autonomous use in high-stakes domains such as medical, legal, financial, or safety-critical decision-making.

## Acknowledgements

O2-9B-Preview is initialized from Qwen3.5-9B. We thank the teams who built and operated the office, customer-service, software-development, verification, and evaluation runtimes that made this model possible.