메뉴
BL
The Decoder • 38일 전

AI 에이전트용 검색 API 품질·비용·속도 벤치마크 공개

IMP
6/10
핵심 요약

Artificial Analysis가 AI 에이전트 관점에서 검색 API 제공자의 품질, 비용, 속도를 측정하는 'Search Index' 벤치마크를 공개했습니다. Parallel, Exa, Firecrawl, Tavily 등 7개 제공자를 동일한 조건에서 비교한 결과, 검색 품질이 좋을수록 토큰 사용량이 줄어 총 비용도 낮아지는 것으로 나타났습니다. Parallel, Firecrawl 등이 비용 대비 성능에서 최적의 조합을 보였습니다.

번역된 본문

새로운 벤치마크, AI 에이전트용 검색 API를 품질·비용·속도 기준으로 순위 매기다

Artificial Analysis가 AI 에이전트에서 검색 API 제공자가 얼마나 잘 작동하는지 품질, 비용, 속도를 기준으로 측정하는 'Search Index' 벤치마크를 출시했다. 초기 대상에는 Parallel, Exa, Firecrawl, You.com, Tavily, Keenable, Brave가 포함된다.

각 제공자는 동일한 모델(GPT-5.6 Luna)로 표준화된 에이전트 환경에서 테스트된다. 변경되는 것은 검색 제공자뿐이다. 에이전트는 Artificial Analysis의 오픈소스 프레임워크인 Stirrup 위에서 작동하며, 태스크당 25회 실행해 웹 검색과 페이지 로드를 수행한다.

이 지수는 동일한 가중치를 가진 세 개의 벤치마크를 결합한다. DeepSearchQA는 각각 여러 검색 쿼리가 필요한 900개의 리서치 질문으로 구성된다. BrowseComp 하위 집합은 다단계 탐색이 필요한 200개의 찾기 어려운 사실을 테스트한다. AA-Omniscience는 6개 지식 영역에 걸친 600개의 질문을 다룬다. 모델이 도구 없이 스스로 답하는 툴 프리 베이스라인이 비교 기준점을 제공한다.

더 나은 검색 품질은 총 비용도 낮춘다. 처음부터 좋은 결과를 받으면 모델이 사용하는 토큰이 줄어들기 때문이다. Parallel Search(advanced)는 Basic 버전에 비해 토큰 사용량이 40% 이상 감소했다. 태스크당 검색 비용은 증가하지만 총 비용은 오히려 더 낮다($0.084 대 $0.11).

쿼리당 날것의 속도가 항상 전체 결과의 빠름을 의미하지는 않는다. Parallel Search(turbo)는 쿼리당 가장 짧은 응답 시간(0.51초, Basic은 1.03초)을 기록했지만, 낮은 품질(67 대 73) 때문에 에이전트가 더 많은 실행을 거쳐야 한다. 결과적으로 태스크당 총 소요 시간은 거의 비슷해진다.

Artificial Analysis에 따르면 Parallel, Firecrawl, Parallel(turbo)이 비용과 성능의 최적 조합을 보였다. 다른 제공자도 벤치마크 참여를 신청할 수 있으며, 전체 방법론은 공개되어 있다.

원문 보기
원문 보기 (영어)
New benchmark ranks search APIs for AI agents on quality, cost, and speed Matthias Bastian View the LinkedIn Profile of Matthias Bastian Aug 18, 2026 Nano Banana Pro prompted by THE DECODER Ask about this article… Search Artificial Analysis has released the "Search Index," a benchmark that measures how well search API providers work for AI agents across quality, cost, and speed. The initial lineup includes Parallel, Exa, Firecrawl, You.com, Tavily, Keenable, and Brave. Each one is tested with the same model (GPT-5.6 Luna) in a standardized agent setup. Only the search provider changes. The agent runs on Stirrup , an open-source framework from Artificial Analysis, and gets 25 runs per task to search and pull up web pages. The index combines three equally weighted benchmarks. DeepSearchQA has 900 research questions, each requiring multiple search queries. A BrowseComp subset tests 200 hard-to-find facts that need multi-step browsing. AA-Omniscience covers 600 questions across six knowledge domains. A tool-free baseline, where the model answers on its own, provides the comparison point. Ad Better search quality also lowers total costs. The model uses fewer tokens when it gets good results up front. With Parallel Search (advanced), token use drops by over 40 percent compared to the Basic version. Per-task search costs go up, but total cost comes in lower ($0.084 vs. $0.11). Ad Raw speed per query doesn't always mean faster results overall. Parallel Search (turbo) clocks the shortest response time per query (0.51 seconds vs. 1.03 seconds for Basic), but its lower quality (67 vs. 73) forces the agent to run more passes. Total time per task winds up about the same. Artificial Analysis says Parallel, Firecrawl, and Parallel (turbo) hit the best mix of cost and performance. Other providers can apply to join the benchmark. The full methodology is public. Ad AI News Without the Hype – Curated by Humans Subscribe to THE DECODER for ad-free reading, a weekly AI newsletter, our exclusive "AI Radar" frontier report six times a year, full archive access, and access to our comment section. Subscribe now Source: via X