Omni-Decision

Evidence-Ledger Planning for Omni-Modal Agents

Ming Ma1,2, Yi Zhu3,*, Yiran Zhong3,*, Feida Zhu3, Yuhao Wang4,
Junhan Shi5, Lingrui Mei6, Tianming Yang1, Steven Hoi3

Affiliations

1 Institute of Neuroscience, Chinese Academy of Sciences
2 University of Chinese Academy of Sciences · 3 Tongyi Lab, Alibaba Group
4 Shanghai Jiao Tong University · 5 Tsinghua University
6 Institute of Computing Technology
* Corresponding authors: Yi Zhu and Yiran Zhong

Keep track of what is known, what is missing,
and what to investigate next.

From question to evidence.

Interactive walkthrough
ExecutionReady to explore

Every answer starts
with an evidence need.

Follow the decisions, one step at a time.

    Illustrative walkthrough adapted from the paper’s travel-video example. Observation excerpts and ledger states are edited for presentation. Select a step to inspect its evidence and ledger.

    Plan from the evidence.

    Omni-Decision maintains an explicit evidence ledger across video, audio, web pages, and computation. The planner selects the next action from the current evidence needs. A critic checks each observation, and the ledger keeps the useful evidence, remaining gaps, and conflicts. The agent answers when the evidence is sufficient.

    Results on OmniGAIA

    Accuracy (%)

    Overall, difficulty, and category breakdown on 360 questions. Best and second-best scores are marked in each column.

    OmniGAIA overall, difficulty and category accuracy in percent.
    MethodOverallDifficultyCategory
    EasyMediumHardGeo.Tech.Hist.Fin.SportsArtMovieSci.Food
    End-to-end models
    Qwen3-Omni-30B* 13.30 19.70 10.60 9.00 8.7014.3011.9028.0010.8013.909.1015.4022.20
    Qwen3.5-Omni-Flash* 33.90 — — — —————————
    Qwen3.5-Omni-Plus* 57.20 — — — —————————
    Gemini-2.5-Pro* 30.80 41.80 26.90 21.80 23.2028.6032.8020.0032.4041.7042.4026.9033.30
    Gemini-3-Flash* 51.70 67.20 46.90 37.20 50.7057.1044.8048.0059.5055.6054.6038.5061.10
    Gemini-3-Pro* 62.50 78.70 61.90 38.50 65.2059.2062.1072.0078.4052.8048.5042.3088.90
    Gemini-3.1-Pro† 79.44 86.07 79.38 69.23 85.5179.5980.3084.0083.7869.4463.6480.7783.33
    Agent systems
    Minimal agent† 5.56 9.02 3.12 5.13 1.4510.2010.610.002.702.786.063.8511.11
    OmniGAIA base† 18.33 26.23 17.50 7.69 5.8028.5728.7916.0013.5113.8915.1526.9211.11
    OmniAtlas-Qwen3-30B* 20.80 31.10 18.80 9.00 10.1030.6029.9032.0018.9016.7012.1011.5027.80
    OmniAgent† 25.51 31.59 24.59 17.89 20.3028.5828.8622.4116.6529.5730.5626.8928.01
    Orchestra-o1-GPT-5* 72.80 80.30 75.00 56.40 72.5069.4075.8064.0083.8063.9069.7073.1083.30
    Sandboxed agent* 75.00 82.00 75.00 64.10 —————————
    Omni-Decision (ours)† 81.39 90.98 78.75 71.79 82.6185.7181.8280.0086.4972.2278.7980.7777.78

    * Publicly reported results. † Our measurements. — indicates an unreported result. See the paper for evaluation details and sources.

    Citation

    BibTeX
    @article{ma2026omnidecision,
      title   = {Omni-Decision: Evidence-Ledger Planning for Omni-Modal Agents},
      author  = {Ma, Ming and Zhu, Yi and Zhong, Yiran and Zhu, Feida and Wang, Yuhao and
                 Shi, Junhan and Mei, Lingrui and Yang, Tianming and Hoi, Steven},
      journal = {arXiv preprint arXiv:2607.11433},
      year    = {2026}
    }