Model ML expands GPT-5.6 Sol for editable finance deliverables

Model ML has expanded use of GPT-5.6 Sol in production finance workflows after testing the model across its Composite evaluation benchmark. The company said GPT-5.6 Sol used 36% fewer tokens per Excel workbook than Opus 5, completed PowerPoint workflows in 100% of test cases, and achieved a 43.3% professional-readiness rate for decks, compared with 26.7% for Opus 5.
The finance automation company uses agents to take assignments from an initial brief through research, analysis and delivery of editable PowerPoint presentations or Excel workbooks. Its core agent plans work, chooses tools, reconciles evidence and runs calculations, while Model ML’s document tooling creates native files with traceable sources.
Tests focus on review-ready files
Model ML’s Composite evaluation follows finance assignments through research and calculations to a deck or spreadsheet, then checks numbers, sources, formulas, structure and, for presentations, visual quality. The company said its PowerPoint benchmark covered hundreds of generated decks and used a detailed scoring rubric.
In the presentation results, GPT-5.6 Sol recorded a 59.9% readiness-gated overall deck-quality score, versus 56.7% for Opus 5. It used 1.10 million tokens per deck, compared with 1.40 million for Fable 5, a reduction of about 21%. GPT-5.6 Sol also had higher scores than Opus 5 for deck quality, brief adherence, hierarchy and design consistency, although Opus 5 scored higher on the aggregate visual judge.
For Excel, GPT-5.6 Sol used 2.44 million tokens per workbook, against 3.83 million for Opus 5. Both models achieved 100% for expected outputs located and for the workbook contract, which covers structure, errors and placeholders. GPT-5.6 Sol took seven minutes per workbook in the reported evaluation, while Opus 5 took 7.5 minutes.
Agent tooling and Microsoft Office workflows
Model ML describes its product as surface-agnostic: users can begin an assignment in email or the Model ML application and continue through Microsoft Office plug-ins without restating the task. The finished documents remain editable, including PowerPoint graphs and tables, rather than being returned as flattened output.
This model of document work aligns with Microsoft 365 Copilot document and data workflows as organisations evaluate how AI assistance changes work with documents and data in Microsoft 365. Model ML said an agent can start from a client template or blank workbook, collect data, build formulas across multiple tabs and apply finance-specific formatting.
The company said a bespoke tearsheet at one global asset manager fell from about one hour of analyst work to about five minutes. In another workflow, its agents processed virtual data rooms with more than 100,000 rows and hundreds of files in one pass.
Practical implication for finance teams
Model ML’s reported results underline that finance teams should assess AI document workflows on auditable deliverables, not just generated prose or visual polish. A practical deployment standard is whether reviewers can trace numbers to sources, recalculate the workbook and edit the slide before a file is shared.

