Perplexity Sonar Retires Sept 27: Agent API Migration Guide | byteiota.
The page describes NEEDLE, an open-source live search benchmark by Keenable that regenerates its query sets hourly for news, daily for finance, scholar, legal, and rare-entity queries, using 15 search APIs under one protocol, with an ultimate oracle engine and reported performance/latency metrics; it notes NEEDLE is…
Keenable AI has open-sourced NEEDLE, a live benchmark for web search APIs that rebuilds its query set from fresh public sources. News queries regenerated hourly; finance, scholar, legal and rare-entity queries regenerated daily. Benchmarks evaluate multiple verticals (News, Everyday, Expert, Deep-tail, Legal) with…
Prices: Perplexity Search API at $5 per 1,000 requests; Exa at $7 per 1,000 requests; Tavily Basic pay-as-you-go at $8 per 1,000 requests. Publication confirms live pricing verified August 29, 2026. No Tavily row exposed in public table. Perplexity Medium on Artificial Analysis Search Index at 80; Exa Auto at 74;…
Google Gemini 3.7 Flash is offered in paid AI Mode for Google AI Pro and Ultra subscribers at $0.75 per million input tokens and $3.75 per million output tokens from now until Dec 31, 2026; from Jan 1, 2027 prices rise to $1.50 and $7.50 per million tokens, respectively. It also marks a global rollout in Search AI…
— The Perplexity Search API has achieved top rankings on the Artificial Analysis Search Index, securing the first three positions. Its medium setting reportedly outperformed previous leaders by five points and advanced the quality-cost frontier to approximately $0.091 per task. The API offers three context settings:…
Not a material, newly published change to a web-search / retrieval / deep-research system. The page discusses latency measurements, provider comparisons, and costs but does not announce a new system, product launch, pricing change, or API modification for a specific web-search/deep-research system with verifiable,…
The page reports that Perplexity’s medium Search API scored 80 on the Artificial Analysis Search Index benchmark (launch Aug 18), beating Parallel Search (75) and Exa Search (74). The index combines DeepSearchQA, BrowseComp, and AA-Omniscience (1,700 questions). It evaluated APIs against a fixed GPT-5.6 Luna model.…
The page documents the release of Dynamic Highlights, a research preview that selects relevant tokens across multiple documents for search queries, reducing token count and enabling a new retrieval paradigm. The page states a 12k character budget with a 40% average token efficiency gain and up to ~95% potential token…
(page discusses pricing and rankings: Parallel Turbo/Fast at $10 per 10,000 searches; Exa at $7 per 1,000 searches; Firecrawl starts at $99; Keenable, Perplexity, Particle Radar, Tavily, Brave, etc. with various tiers and claims).
The page documents migration from the Assistants API to the Responses API, noting the Assistants API sunset on 2026-08-26 and detailing new constructs (Prompts, Conversations, Responses) and migration steps. No exact model IDs, tiers, prices, or endpoint changes are provided on this page beyond the sunsetting and…
The page documents OpenWebSearch, a unified search API gateway by Interfaze (JigsawStack) that routes requests across multiple providers (Exa, Brave, Perplexity, Parallel, Valyu, and more) using one API key and one request format, with automatic provider fallback and a normalized response shape (title, url, snippet,…
Accel-backed Keenable is indexing the web for AI agents | TechCrunch.
Fundraising: Keenable announces $26,000,000 in funding led by Accel; no explicit change to a live API or public product pricing documented in this page. Product launch details include two products: a single index powering over 100 billion documents with sub-250ms latency (p95) and Time Machine feature (query_time)…
The page documents a material change: AWS Bedrock AgentCore Web Search now supports domain allowlists up to 100 domains, exclude lists of 100 domains, published-date range filtering using ISO-8601 timestamps, and expansion to Europe and Asia Pacific, at a price of $7 per 1,000 queries, with no separate API key or…
reports OpenAI cutting frontier GPT-5.6 Sol model developer pricing by more than 20% for the next three months; GPT-5.6 Sol now $4 per 1M input tokens and $20 per 1M output tokens (previously $5/$30).
GPT-5.6 Sol API pricing reduced by over 20% for the next 3 months: Sol $4 per million input tokens, $20 per million output tokens; Terra and Luna prices also reduced; promo valid at least through 2026-11-21; applies to API, ChatGPT Work, Codex credits; Pro/Plus/Business unchanged.
(as cited in the thread)
The page cites specific, newly published pricing details for Perplexity Deep Research: Free users get 5 Deep Research runs/day; Pro subscribers at $20/month or $200/year get 20 runs/day; ~600 reports/month implied by 30-day month; Sonar Deep Research API priced at $3.50 per million tokens (blended $2 input, $8…
Gemini 3.7 Flash (gemini-3.7-flash) generally available on the Gemini API on 13 August 2026, described as the most intelligent workhorse model yet for coding and agents, with substantial gains in software engineering, web development and agentic workflows, offered at an introductory price through 31 December 2026.
Claude vs Perplexity: 5x Context Window Gap [2026]. (example placeholder)
Gemini Deep Research agent - Google AI for Developers.
OpenAI’s Chat Completion API now includes models with built-in web search functionality and the Responses API adds built-in tool calls for web and file search, enabling real-time data retrieval and reducing need for external tools. Perplexity Sonar models (web-enabled Llama 3.3 and DeepSeek R1) are available via the…
The page states that Microsoft retired the Bing Search APIs on 2025-08-11 and redirected developers to Grounding with Bing Search inside Azure AI Agents, with higher costs and limited scope, contributing to a shift toward AI-native web search APIs (including Parallel, Exa, Brave) in late 2025 and 2026.
Artificial Analysis tested Perplexity Search API and found medium, high, and low context settings to achieve the top three scores (80, 79, 77 respectively) in their Artificial Analysis Search Index, with Perplexity charging $5 per 1,000 Search requests and showing favorable quality/price trade-offs across context…
Google made Gemini 3.8 Flash available as a new stable Gemini API model aimed at autonomous agents and long-horizon workflows. Its documented capabilities include Search grounding, File Search, URL Context, structured outputs, function calling, computer use in preview, and a 1,048,576-token input limit, materially…
OpenAI announced that eligible ChatGPT for Clinicians users can search Healthcare Public Data across nine public healthcare sources for medical research, clinical trials, medication information, Medicare data, and provider records. The read-only healthcare search capability is also available through healthcare…
Keenable AI released NEEDLE, an open-source live-search benchmark that rebuilds its query set every hour and evaluates 15 search APIs for agent use. The reported results materially distinguish providers on freshness-sensitive search, including Exa at 0.910, Perplexity at 0.871, Keenable at 0.774, and Tavily at 0.310…
You.com launched Web Search with Highlights, an extraction_mode that returns query-focused verbatim passages while preserving numbers, dates, tables, and formatting. You.com reports 95.17% SimpleQA accuracy, a 16.5% relative lift at the same $5 CPM, 695 ms p50 latency, and independent Braintrust testing showing up to…
Parallel published updated Web Search API quality benchmarks covering BrowseComp, FRAMES, SimpleQA, Coding, HLE, SealQA, and WebWalker. The page reports Parallel Turbo at 51% accuracy and 216 ms p50 latency on BrowseComp, compared with Exa Instant at 33.7% and 361 ms, Tavily Ultra Fast at 19.3% and 357 ms, Brave…
keenable.ai (read 2026-08-26): Agent Builder tier $4 / 1000 requests; Frontier tier $1 / 1000 requests at 100 RPS+; <250 ms p95 (US East); 100B+ documents.
Research guide (read 2026-08-26): research_effort frontier at $1,200 per 1k calls, background=true required (sync → 422), 30 s–12,000 s with p50 300 s. Other tiers: lite $12, standard $50, deep $100, exhaustive $450 per 1k. Not wired in the arena.
OpenAI's deprecations page (read 2026-08-26) has no deep-research entry. The 2026-12-11 shutdown announced 2026-06-11 covers o3-2025-04-16, o3-pro-2025-06-10 and gpt-5-*-2025-08-07 snapshots → gpt-5.6-sol / terra / luna. gpt-5.5-pro is live at $30 / $180 per 1M. research.site's earlier “o3-deep-research retires 2026-12-11” item was wrong and is withdrawn.
Changelog 2026-08-24: citations and confidence scores are now emitted per list element (separate FieldBasis entries) instead of once per field, without the beta header.
OpenAI changelog 2026-08-21: gpt-5.6-sol pricing updated to $4.00 / 1M input and $20.00 / 1M output.
Changelog 2026-08-21: Search API mode "fast" at $1 per 1,000 requests with ~700 ms average latency; Turbo mode (2026-07-13) is $1 per 1,000 with p50 250 ms; basic / advanced stay $5 per 1,000. Task API processor prices unchanged ($5 lite → $2,400 ultra8x per 1k runs).
Changelog 2026-07-21: Responses API for latency-sensitive use, 5–60 s response times, structured outputs and full citations; synchronous, alongside the async Task API.
Claude Code vs Codex CLI under three search backends (builtin, Exa MCP, Valyu MCP) on a frozen 30-task LiveBrowseComp subset. claude-exa wins at 60%; the search backend moves accuracy up to 33 points, the agent at most 7.
Five deep-research APIs plus a Perplexity calibration anchor on a frozen 30-task ResearcherBench subset. Coverage is a three-way race; faithfulness splits the field in two; parallel has the best overall profile.
Nine deep research APIs on a frozen 20-task subset of DeepResearch-Bench-II. Exa effort tiers against each other, then one flagship tier per provider.