Observed Signal · Apr 22, 2026 · Technical Release · Source: OpenAI Blog · Impact: 4/5 · Sentiment: Positive
WebSockets speed up Responses API agentic workflows
OpenAI announced a WebSocket mode for its Responses API to reduce API overhead for agentic workflows. By keeping a persistent connection and caching previous response state, the Responses API can avoid repeated tokenization and validation work, overlap pipeline stages, and only process new input. Combined with caching, fewer network hops, and faster safety checks, OpenAI reports end-to-end agent loop speedups of about 40%, enabling GPT‑5.3‑Codex‑Spark to run at roughly 1,000 tokens-per-second (TPS) with bursts to 4,000 TPS. An alpha with coding-focused partners (including Vercel, Cline, and Cursor) reported latency improvements; the feature preserves the familiar response.create call shape via a previous_response_id mechanism and an in-memory connection-scoped cache. The blog post is dated April 22, 2026 and authored by Brian Yu and Ashwin Nathan.
Technical release from a major AI platform (OpenAI) that materially reduces API latency for LLM agentic workflows, enabling higher inference throughput and affecting builders and MarTech applications that integrate fast agents.
Track Vercel Signals & Market Shifts in Real-Time
Polaris7 autonomous intelligence agents track regulatory filings, primary sources, executive changes, and deal flow 24/7. Create your free Explorer workspace to monitor these entities.
Key Takeaways & Evidence Grounding
- OpenAI launched a WebSocket mode for the Responses API to provide persistent connections and in-memory cached response state.
- Agentic workflows using WebSocket mode achieved up to ~40% end-to-end speed improvements in latency compared with sequential HTTP requests.
- GPT‑5.3‑Codex‑Spark achieved approximately 1,000 tokens-per-second (TPS) in production with bursts up to 4,000 TPS after the launch.
- Optimizations included caching rendered tokens and model configuration, eliminating intermediate network hops, and faster safety classifier processing.
- Alpha integrations with partners (Vercel, Cline, Cursor) reported latency decreases (examples: up to 40%, 39%, and 30% respectively).
Connected Companies & Entities
2 Entities mappedOntology Mapping & Concepts
Related Market Signals & Shifts
Recent verified developments and strategic activity across this market segment.
OpenAI Unveils Lightning-Fast GPT-5.3 Codex-Spark Model
OpenAI announced a research preview of GPT‑5.3‑Codex‑Spark, a smaller, ultra-fast variant of GPT‑5.3‑Codex optimized for real‑time coding. Launched February 12, 2026, Codex‑Spark is served on low‑latency Cerebras hardware and delivers over 1,000 tokens per second with a 128k text-only context window. Initially available to ChatGPT Pro users as a research preview and to a small set of API design partners, Codex‑Spark emphasizes minimal, targeted edits for interactive workflows. OpenAI also rolled out end-to-end latency improvements (persistent WebSocket support and Responses API optimizations) that reduce roundtrip and per-token overheads and speed time-to-first-token. The model is evaluated under OpenAI’s safety processes and is text-only at launch; broader capabilities and access will expand over time.
OpenAI Unveils Lightning-Fast Codex with New Cerebras Chip
OpenAI announced a lightweight, low-latency version of its coding agent Codex called GPT-5.3-Codex-Spark, designed for faster inference and rapid prototyping. Spark is intended for real-time collaboration and shorter development tasks, and is available as a research preview to ChatGPT Pro users in the Codex app. To deliver the lower latency, OpenAI is running Spark on Cerebras’ Wafer Scale Engine 3 (WSE-3) as part of a recently disclosed multi-year compute agreement between the two companies reportedly worth over $10 billion. Cerebras’ WSE-3 is described as its third-generation waferscale megachip with roughly 4 trillion transistors. Cerebras also recently raised $1 billion in funding at a reported $23 billion valuation. OpenAI framed Spark as the first milestone in deeper hardware integration with Cerebras to speed model responses.
Two API Settings Tripled ARC‑AGI‑3 Scores
OpenAI reports that enabling two Responses API settings—retained reasoning and compaction—substantially improved agent performance on the ARC‑AGI‑3 2D puzzle benchmark. Using their Responses API harness (which retains private reasoning messages across turns and compacts long histories instead of rolling truncation), GPT‑5.6 Sol's score on the ARC‑AGI‑3 public set rose from 13.3% with the official harness to 38.3%, roughly a 3x improvement, while output tokens fell by about 6x. The post explains that the official ARC harness discarded private reasoning and used rolling truncation, which prevented models from carrying forward internal thoughts and older actions. OpenAI recommends using retained reasoning and compaction in evaluations to better match production deployments like ChatGPT and Codex.
Track Real-Time Market Signals & Shifts
Set up custom watchlists to receive automated, evidence-grounded executive digests whenever material signals or shifts occur across your tracked landscape.
