Omni-Decision
Evidence-Ledger Planning for Omni-Modal Agents
Affiliations
1 Institute of Neuroscience, Chinese Academy of Sciences
2 University of Chinese Academy of Sciences · 3 Tongyi Lab, Alibaba Group
4 Shanghai Jiao Tong University · 5 Tsinghua University
6 Institute of Computing Technology
* Corresponding authors: Yi Zhu and Yiran Zhong
Keep track of what is known, what is missing,
and what to investigate next.
From question to evidence.
Interactive walkthroughEvery answer starts
with an evidence need.
Follow the decisions, one step at a time.
Illustrative walkthrough adapted from the paper’s travel-video example. Observation excerpts and ledger states are edited for presentation. Select a step to inspect its evidence and ledger.
Plan from the evidence.
Omni-Decision maintains an explicit evidence ledger across video, audio, web pages, and computation. The planner selects the next action from the current evidence needs. A critic checks each observation, and the ledger keeps the useful evidence, remaining gaps, and conflicts. The agent answers when the evidence is sufficient.
Results on OmniGAIA
Accuracy (%)Overall, difficulty, and category breakdown on 360 questions. Best and second-best scores are marked in each column.
| Method | Overall | Difficulty | Category | ||||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| Easy | Medium | Hard | Geo. | Tech. | Hist. | Fin. | Sports | Art | Movie | Sci. | Food | ||
| End-to-end models | |||||||||||||
| Qwen3-Omni-30B* | 13.30 | 19.70 | 10.60 | 9.00 | 8.70 | 14.30 | 11.90 | 28.00 | 10.80 | 13.90 | 9.10 | 15.40 | 22.20 |
| Qwen3.5-Omni-Flash* | 33.90 | — | — | — | — | — | — | — | — | — | — | — | — |
| Qwen3.5-Omni-Plus* | 57.20 | — | — | — | — | — | — | — | — | — | — | — | — |
| Gemini-2.5-Pro* | 30.80 | 41.80 | 26.90 | 21.80 | 23.20 | 28.60 | 32.80 | 20.00 | 32.40 | 41.70 | 42.40 | 26.90 | 33.30 |
| Gemini-3-Flash* | 51.70 | 67.20 | 46.90 | 37.20 | 50.70 | 57.10 | 44.80 | 48.00 | 59.50 | 55.60 | 54.60 | 38.50 | 61.10 |
| Gemini-3-Pro* | 62.50 | 78.70 | 61.90 | 38.50 | 65.20 | 59.20 | 62.10 | 72.00 | 78.40 | 52.80 | 48.50 | 42.30 | 88.90 |
| Gemini-3.1-Pro† | 79.44 | 86.07 | 79.38 | 69.23 | 85.51 | 79.59 | 80.30 | 84.00 | 83.78 | 69.44 | 63.64 | 80.77 | 83.33 |
| Agent systems | |||||||||||||
| Minimal agent† | 5.56 | 9.02 | 3.12 | 5.13 | 1.45 | 10.20 | 10.61 | 0.00 | 2.70 | 2.78 | 6.06 | 3.85 | 11.11 |
| OmniGAIA base† | 18.33 | 26.23 | 17.50 | 7.69 | 5.80 | 28.57 | 28.79 | 16.00 | 13.51 | 13.89 | 15.15 | 26.92 | 11.11 |
| OmniAtlas-Qwen3-30B* | 20.80 | 31.10 | 18.80 | 9.00 | 10.10 | 30.60 | 29.90 | 32.00 | 18.90 | 16.70 | 12.10 | 11.50 | 27.80 |
| OmniAgent† | 25.51 | 31.59 | 24.59 | 17.89 | 20.30 | 28.58 | 28.86 | 22.41 | 16.65 | 29.57 | 30.56 | 26.89 | 28.01 |
| Orchestra-o1-GPT-5* | 72.80 | 80.30 | 75.00 | 56.40 | 72.50 | 69.40 | 75.80 | 64.00 | 83.80 | 63.90 | 69.70 | 73.10 | 83.30 |
| Sandboxed agent* | 75.00 | 82.00 | 75.00 | 64.10 | — | — | — | — | — | — | — | — | — |
| Omni-Decision (ours)† | 81.39 | 90.98 | 78.75 | 71.79 | 82.61 | 85.71 | 81.82 | 80.00 | 86.49 | 72.22 | 78.79 | 80.77 | 77.78 |
* Publicly reported results. † Our measurements. — indicates an unreported result. See the paper for evaluation details and sources.
Citation
BibTeX@article{ma2026omnidecision,
title = {Omni-Decision: Evidence-Ledger Planning for Omni-Modal Agents},
author = {Ma, Ming and Zhu, Yi and Zhong, Yiran and Zhu, Feida and Wang, Yuhao and
Shi, Junhan and Mei, Lingrui and Yang, Tianming and Hoi, Steven},
journal = {arXiv preprint arXiv:2607.11433},
year = {2026}
}