
Anthropic’s tests show AI agents can clash, collude and coordinate in ways that complicate safety testing.
Artificial Analysis’ new Search Index compares seven search API providers for AI-agent workflows, with Parallel, Exa, and Firecrawl leading the initial quality rankings.

Artificial Analysis has released the Search Index, a benchmark for evaluating search API providers used by AI agents. The index rates providers across quality, cost, and speed, giving teams a clearer way to compare search tools for agentic workflows. The initial test group includes Parallel, Exa, Firecrawl, You.com, Tavily, Keenable, and Brave.

Each provider is tested with the same model, GPT-5.6 Luna, in a standardized agent setup where only the search provider changes. The agent runs on Stirrup, Artificial Analysis’ open-source framework, and gets 25 runs per task to search and retrieve web pages. The index combines three equally weighted benchmarks: DeepSearchQA with 900 research questions, a BrowseComp subset with 200 hard-to-find facts, and AA-Omniscience with 600 questions across six knowledge domains.
The tool-free baseline scored 33 points, while search-enabled scores ranged from 65 to 75. Parallel, Exa, and Firecrawl led the initial ranking with scores of 75, 74, and 73, respectively. The results show that giving an AI agent access to search can substantially improve task performance compared with relying on the model alone.
The benchmark also highlights a practical tradeoff: better search quality can reduce total costs because the model uses fewer tokens when useful results arrive earlier. In the article’s example, Parallel Search advanced cut token use by over 40 percent compared with the Basic version. Even though per-task search costs increased, total cost was lower at $0.084 versus $0.11.
Raw query speed does not always translate into faster task completion. Parallel Search turbo had the shortest per-query response time at 0.51 seconds versus 1.03 seconds for Basic, but its lower quality score meant the agent needed more passes. Artificial Analysis says Parallel, Firecrawl, and Parallel turbo offer the strongest mix of cost and performance in the current results.

Anthropic’s tests show AI agents can clash, collude and coordinate in ways that complicate safety testing.

An open-weight Meta model points to local personal AI agents — and a divide between access and ownership.

Kitesurf is a cloud-hosted browser for AI agents, built to help developers automate browser-based tasks more efficiently.

A Claude Code usage log suggests agentic AI can consume far more electricity than simple chat queries.