— OpenAI announced Dots AI agents built on GPT-6 Astra and the GPT-6.1 Sol model at DevDay, with pricing for GPT-6.1 Sol in API: $2 per 1M input tokens and $10 per 1M output tokens, cached input $0.10 per 1M tokens; Sol features include long-running tasks, efficiency gains, and availability via API, ChatGPT Work, and…
Photon is Perplexity's in-house Rust-based retrieval and ranking engine that replaces a forked open-source engine; production traffic now served by Photon with Fast Search mode in the Perplexity Search API. Production p99 latency for Photon’s stages fell from ~800 ms to ~65 ms; Photon runs on ~20% fewer serving…
Keenable + cognee: Persistent Memory for Web Retrieval.
Gemini API shuts down the May Antigravity agent on 5 October, with renamed tools — TrustList.
– The article describes Perplexity Agent API as a materially different product from Sonar, including a new endpoint /v1/agent (with alias /v1/responses), model-agnostic access to multiple frontier providers, built-in tools (web search with advanced filters, fetch_url, finance_search), a redefined pricing model with…
URL: Perplexity’s Sonar Chat Completions endpoint retires on September 27, 2026, migrating to the Agent API at with explicit web_search tooling requirements and response structure changes; OpenAI-compatible SDKs point to /v1 and /v1/responses; embedded citations and citation data move inside the agent output via…
(URL cited as source) details: Perplexity retires the Sonar chat completions API on September 27, 2026. The Sonar endpoint moves to Perplexity's Agent API on September 25 with unchanged request shape but new transport; perplexity/sonar-pro and perplexity/sonar-reasoning-pro stop being routable on September 27. The…
Perplexity’s Sonar API is retiring on September 27, 2026; Sonar calls will stop working; replacement is the Agent API with new presets mapping Sonar -> fast, Sonar Pro -> low, Sonar Reasoning Pro -> medium, Sonar Deep Research -> high; Agent API supports multi-step reasoning, tools (fetch_url, finance_search), and…
Perplexity retires the Sonar API on September 27, 2026; Appsmith shuts down its managed Appsmith AI plugin on September 30, 2026; migration to Perplexity Agent API with new endpoints and pricing/rate-limits, and Appsmith OpenAI/Anthropic/Google AI datasources migration path; end of Sonar support dates and deprecation…
Parallel Web Systems provides a web search API with four freshness/latency modes (Turbo, Fast, Basic, Advanced) and an MCP server at parallel.ai/mcp exposing web_search and web_fetch tools without an API key; with API key, rate limits and tool call allowances apply; pricing noted as 25,00 within MCP context.
— Exa releases Agent Ultra, the highest effort level of its Exa Agent API, deployable as a hosted API (effort: "ultra"), not open weights or self-hostable, with performance benchmarks against Opus 5.5, GPT-6 Astra, and Perplexity Agent and pricing details including default $20 per run and max cost between $1-$100.
OpenAI GPT-6 Sol: $2 per 1M input tokens, $10 per 1M output tokens; GPT-6 Luna: $0.10 per 1M input tokens, $0.50 per 1M output tokens. Comparative benchmarks show Sol at 33.2% vs Opus 5 at 26.9% and cost per task significantly lower; Sol and Luna released Sep 22; pricing cited as a 50% cut against GPT-5.6 pricing.
documents that on September 25, 2026, perplexity/sonar moves to Perplexity's Agent API with the same model id and lower price; perplexity/sonar-pro and perplexity/sonar-reasoning-pro stop being routable on September 27, 2026; full details and timing are provided in the changelog page.
Exa Agent Ultra is introduced as the highest effort level for Exa Agent, coordinating thousands of sources, delivering frontier-level results at a fraction of the cost; benchmarks claim Ultra outperforms Opus 5.5, GPT-6 Astra, and Perplexity Agent on their respective maximum effort settings; pricing notes indicate…
- Photon: Building a Retrieval and Ranking Engine From Scratch announces Photon, an in-house retrieval and ranking engine, its performance metrics (single-search-call latency 160 ms p50, 230 ms p95 for fast preset; latency improvements from 800 ms to 65 ms p99; 2.5x data per document; 68% cost reduction; six agentic…
The page documents material changes: the Perplexity Sonar functionality now uses the Agent API at POST with model perplexity/sonar and web_search tool; legacy sonar endpoints remain working until Sep 27, 2026; pricing for Agent API uses $0.25 per million input tokens, $2.50 per million output tokens, and $0.0025 per…
Perplexity launched Fast Search, a mode of its Search API, with a median latency of 160 ms and 95% within 230 ms, priced at $1 per 1,000 requests, using Photon (Rust-based retrieval and ranking engine). The mode supports up to 20 results per request, requires search_type="fast" in POST /search, and is targeted at AI…
Photon: Building a Retrieval and Ranking Engine From Scratch.
Photon, Perplexity’s in-house retrieval and ranking engine, now powers the production search stack, replacing a third-party open-source engine; Fast Search introduced as a lower-cost option at $1.00 per 1,000 requests, alongside standard web search at $5.00 per 1,000 requests.
Source notes: - Microsoft retires Bing Search APIs on 2025-08-11; replacement is Grounding with Bing Search inside Azure AI Foundry, billed around $35 per 1,000 transactions. - 2026 market shift includes agent-native APIs from Exa (≈$7 per 1,000 with contents), Parallel (funding and API launch), Linkup, Tavily, etc.…
Google’s Gemini 4 Enters Post-Training — and the Three-Way Frontier Race Just Compressed – Forkast.
The page states the knowledge parameter is added to the web search call with the You.com Web Search API at the same $5 CPM, increasing accuracy from 64.5% to 84.2% on VerticalRTK benchmark, with latency ~0.60s baseline and ~0.74s with knowledge. The exact parameter is: "knowledge": "core". Full parameter reference:
announces two new text-to-speech models: Gemini 3.8 Flash TTS and Gemini 3.8 Flash-Lite TTS, with deployment starting today in Google AI Studio, Gemini API, Google Vids, and Gemini Enterprise (API coming soon); describes capabilities, use cases, and partner integrations.
Effort Mode, GPT-6 Astra & Skills Marketplace | Perplexity.
documents a comparative benchmark between Tavily and Apify's RAG Web Browser, including retrieval styles, capability comparisons, and cost metrics for a benchmark involving four queries; no explicit new product launch, pricing change, or API change is announced on this page. The page is a comparative benchmark…
Parallel’s official Search API page now documents the selected fast-search mode as “fast” rather than “turbo” in the previous snapshot. Because search modes are part of the API’s endpoint behavior and system record, this is a material mode-naming or configuration change requiring verification in the provider…
Parallel materially changed its Search API offering: the fast search mode is now labeled “turbo” instead of “basic,” and the displayed cost decreased from $5 to $1 per 1,000 requests. The page continues to describe multiple search modes with different latency/depth tradeoffs, making this a material pricing and…
Openbenchmarks materially updated its independent web-search evaluation board on September 15. The update added Nimble lite and standard configurations across company-news, coding-agent, and multi-turn search benchmarks, plus You.com highlights with and without knowledge=core. The board now reports You highlights at…
Parallel materially updated its Search API offering with four search modes: Turbo at about 200 ms and $1 per 1,000 requests, Fast at about 700 ms and $1 per 1,000, Basic at about 1 second and $5 per 1,000, and Advanced at about 3 seconds and $5 per 1,000. The page also states a 600-requests-per-minute limit and…
Perplexity published a major Advanced Deep Research update. The system now searches more sources, cross-references information, processes uploaded documents, runs calculations through an improved code sandbox, and browses harder-to-reach web sources. It also adds clarifying questions, follow-up questions during…
OpenAI published a Perplexity deployment update stating that Perplexity uses GPT-6 Astra to write communications, modify software, test applications, and monitor production systems. Perplexity also reports that Astra improves programs that search the web and internal information, and that it can be trusted with…
DeepSearch announced a new AI research platform for professionals that combines public-web research, source-linked findings, and contextual follow-up questions in one workspace. The platform is positioned as a source-backed research system that lets users inspect original material, compare accounts, identify evidence…
Openbenchmarks added new quantitative headline results to its independent web-search benchmark: Exa fast leads factual lookup at 99.3% across 300 questions, Exa deep leads hard retrieval at 83.0% across 100 tickets, and Parallel basic leads multi-hop search at 46.5% F1 across 45 questions. The page also clarified…
Openbenchmarks materially updated its independent web-search benchmark. The September 11 update replaced Tavily Fast with Tavily Basic on the search-only hard-retrieval board, evaluated 100 tasks across three repeats, replaced Tavily ultra-fast with Basic and Advanced on the company-news benchmark, and changed…
OpenAI released the Agents API in public beta. The API provides a managed Codex harness for session orchestration, context compaction, recovery, durable multi-turn sessions, progress streaming, custom tools and MCP servers, with either OpenAI-hosted sandboxes or supported external infrastructure. This is a new…
Openbenchmarks published an independent web-search benchmark update covering factual lookup, hard retrieval, and multi-hop search. The page reports a September 9, 2026 re-evaluation of TinyFish after updates to its GA Fetch endpoint, and provides quantitative provider comparisons using accuracy, answer recall,…
Parallel materially expanded its official quality benchmarks with new direct comparisons against Perplexity, Exa Auto, and Tavily across SimpleQA Verified, FRAMES, BrowseComp, and WideSearch. The revised methodology states that every competitor was evaluated at both low-cost and frontier tiers using the same agent,…
Parallel’s official benchmarks page now reports a new September 9, 2026 evaluation round for its Search API and Task API. The update adds fresh quantitative results across SimpleQA Verified, FRAMES, BrowseComp, and WideSearch, comparing Parallel Fast, Turbo, Advanced, and Basic tiers against OpenAI Web Search.…
Fortune reported that OpenAI repeatedly changed GPT-6 Astra’s published evaluation results after launch. Astra’s reported hallucination rate changed from 4.2% to 2% and later returned to 4.2%; its ARC-AGI-3 score changed from 98.6% in a pre-publication draft to 99.99% in the live post; and GPT-5.6 Sol’s ExploitBench…
OpenAI introduced GPT-6 Astra in a limited organizational rollout as a new model for research, coding, computer use, and complex multi-step work. The release adds the GPT-6 Astra model to the Responses and Chat Completions APIs, while the Responses API gains material agent-workflow controls including asynchronous…
Google made Gemini 3.8 Flash available as a new stable Gemini API model aimed at autonomous agents and long-horizon workflows. Its documented capabilities include Search grounding, File Search, URL Context, structured outputs, function calling, computer use in preview, and a 1,048,576-token input limit, materially…
OpenAI announced that eligible ChatGPT for Clinicians users can search Healthcare Public Data across nine public healthcare sources for medical research, clinical trials, medication information, Medicare data, and provider records. The read-only healthcare search capability is also available through healthcare…
You.com launched Web Search with Highlights, an extraction_mode that returns query-focused verbatim passages while preserving numbers, dates, tables, and formatting. You.com reports 95.17% SimpleQA accuracy, a 16.5% relative lift at the same $5 CPM, 695 ms p50 latency, and independent Braintrust testing showing up to…
Parallel published updated Web Search API quality benchmarks covering BrowseComp, FRAMES, SimpleQA, Coding, HLE, SealQA, and WebWalker. The page reports Parallel Turbo at 51% accuracy and 216 ms p50 latency on BrowseComp, compared with Exa Instant at 33.7% and 361 ms, Tavily Ultra Fast at 19.3% and 357 ms, Brave…
Keenable AI released NEEDLE, an open-source live-search benchmark that rebuilds its query set every hour and evaluates 15 search APIs for agent use. The reported results materially distinguish providers on freshness-sensitive search, including Exa at 0.910, Perplexity at 0.871, Keenable at 0.774, and Tavily at 0.310…
OpenAI's deprecations page (read 2026-08-26) has no deep-research entry. The 2026-12-11 shutdown announced 2026-06-11 covers o3-2025-04-16, o3-pro-2025-06-10 and gpt-5-*-2025-08-07 snapshots → gpt-5.6-sol / terra / luna. gpt-5.5-pro is live at $30 / $180 per 1M. research.site's earlier “o3-deep-research retires 2026-12-11” item was wrong and is withdrawn.
Changelog 2026-08-24: citations and confidence scores are now emitted per list element (separate FieldBasis entries) instead of once per field, without the beta header.
OpenAI changelog 2026-08-21: gpt-5.6-sol pricing updated to $4.00 / 1M input and $20.00 / 1M output.
Changelog 2026-08-21: Search API mode "fast" at $1 per 1,000 requests with ~700 ms average latency; Turbo mode (2026-07-13) is $1 per 1,000 with p50 250 ms; basic / advanced stay $5 per 1,000. Task API processor prices unchanged ($5 lite → $2,400 ultra8x per 1k runs).
Changelog 2026-07-21: Responses API for latency-sensitive use, 5–60 s response times, structured outputs and full citations; synchronous, alongside the async Task API.
One GPT-6 Astra agent, five web tool stacks. Exa and built-in web tie at 11/26 on LiveBrowseComp, with Keenable one answer behind. Explore accuracy, run duration, and every public task verdict.
Claude Code vs Codex CLI under three search backends (builtin, Exa MCP, Valyu MCP) on a frozen 30-task LiveBrowseComp subset. claude-exa wins at 60%; the search backend moves accuracy up to 33 points, the agent at most 7.
Five deep-research APIs plus a Perplexity calibration anchor on a frozen 30-task ResearcherBench subset. Coverage is a three-way race; faithfulness splits the field in two; parallel has the best overall profile.
Nine deep research APIs on a frozen 20-task subset of DeepResearch-Bench-II. Exa effort tiers against each other, then one flagship tier per provider.