AI · 1 October 2026
Iris-mini, Iris-pro: AllSpark's open-weight search-agent models
AllSpark's Iris-mini and Iris-pro, built on Alibaba's Qwen models, lead open-weight benchmarks for search-agent tasks and show unexpected strength in tool use and office productivity.
What happened
AllSpark has released two new open-weight language models, Iris-mini and Iris-pro, which the company says are now the strongest open-weight search-agent models in their respective size classes. Built on top of Alibaba's Qwen model family, the pair were designed primarily for search-agent tasks — retrieving, reasoning over and acting on information gathered from external tools — and have topped open-weight benchmarks in that category.
According to reporting from The Decoder, testing also surfaced gains beyond the models' core design brief: both Iris-mini and Iris-pro performed unexpectedly well on tool-use tasks and on office-productivity style workloads, areas the models were not specifically built or tuned for.
Why it matters
Search-agent capability is becoming a proxy for how useful a language model is as an autonomous worker rather than a conversational assistant: the ability to query tools, retrieve accurate information and act on it reliably underpins agentic AI features now being built into customer service, research and enterprise software. A strong open-weight entrant in this category matters because it lowers the barrier for organisations wanting to build or customise agentic AI systems without being locked into closed, proprietary models.
The incidental strength in tool use and office-type tasks is arguably the more interesting signal. It suggests that models optimised for structured, multi-step search and retrieval may generalise better to broader "digital worker" tasks than models trained narrowly for chat or writing — a useful data point for anyone evaluating which open-weight base to standardise on for agentic deployments.
The Renascence take
Benchmark leadership in a single, narrow category is easy to overread. The detail worth sitting with here is the spillover: capability built for one job (search agency) showing up unprompted in another (office productivity). That is a signal about how these models are learning to generalise, not just about who tops a leaderboard this month.
Most organisations still evaluate AI models the way they'd evaluate a vendor brochure — by the headline benchmark. The more useful question is what a model is quietly good at that nobody tested for, because that's where real operational value tends to hide. Teams piloting agentic AI for service or back-office work should be testing open-weight models like Iris-mini and Iris-pro on their own messy, specific tasks before assuming a benchmark win translates into a fit for their workflow.
Sources
This briefing was written by our Newsdesk, synthesising reporting from the outlets below. Follow the links for the original coverage.
FAQ
Questions we get on this topic
More in AI
Stay ahead of CX
Get the signal, not the noise.
The stories shaping customer experience — plus the Journal and Experience Loom — in your inbox.
