Observed Signal · May 27, 2026 · Technical Release · Source: DEV Community · Impact: 3/5 · Sentiment: Positive
Cursor and Fireworks Detail Composer 2 Model
A DEV Community post summarizes a Sequoia Capital podcast featuring Federico Cassano (Cursor) and Dmytro Dzhulgakov (Fireworks) discussing Composer 2, a code-specialized model trained by Cursor on Fireworks' distributed infrastructure. The speakers describe Composer 2's training recipe — continuing pretraining on code using a Kimi 2.5 MoE foundation and large-scale reinforcement learning in Cursor's sandbox — and detail engineering innovations: an asynchronous pipeline to maximize GPU utilization, global distributed inference with incremental 'Delta Sync' weight transfers, GPU kernel fixes and a 'Router Replay' system to address MoE numerical mismatch, and online real-time RL driven by user feedback. They also describe Composer 2's very large effective context handling via self-summarization, and claim inference cost and latency advantages versus larger generalist models.
Describes concrete engineering patterns (asynchronous RL, global incremental weight sync, MoE numerical alignment, realtime online RL) that lower cost and latency for specialized LLMs — relevant to organizations building production ML/LLM infrastructure.
Track Sequoia Capital Signals & Market Shifts in Real-Time
Polaris7 autonomous intelligence agents track regulatory filings, primary sources, executive changes, and deal flow 24/7. Create your free Explorer workspace to monitor these entities.
Key Takeaways & Evidence Grounding
- Sequoia Capital hosted a podcast with Federico Cassano (Cursor) and Dmytro Dzhulgakov (Fireworks) about Composer 2.
- Composer 2 was built by continuing pretraining on a code-heavy corpus from a Kimi 2.5 base model (described as a 1 trillion-parameter MoE with ~30B active parameters).
- Training combined mid-training on code tokens and large-scale reinforcement learning (RL) in Cursor's sandbox/harness to teach tool use and correct code generation.
- Cursor and Fireworks implemented distributed engineering innovations: an asynchronous training/inference pipeline, global inference across four small clusters, and a database-level lossless compression plus incremental weight transfer system ('Delta Sync') that reduced transfer size by about 20x and enabled ~30s–1min global weight syncs.
- They addressed numerical inconsistency in sparse MoE routing by writing custom GPU kernels and using a 'Router Replay' technique to align expert selection between inference and training.
Connected Companies & Entities
4 Entities mappedOntology Mapping & Concepts
Related Market Signals & Shifts
Recent verified developments and strategic activity across this market segment.
Cursor Admits Composer 2 Built on Moonshot AI’s Kimi
Cursor launched a new AI coding model called Composer 2 but faced scrutiny after an X user (Fynn) and code evidence suggested the model was effectively built on top of Moonshot AI’s open-source Kimi 2.5. Cursor VP of developer education Lee Robinson acknowledged Composer 2 “started from an open-source base,” saying roughly one quarter of the compute used for the final model came from that base and the remainder from Cursor’s own training, and that Composer 2 performs differently on benchmarks. The Kimi account and Moonshot‑linked posts said Cursor’s use was part of an authorized commercial partnership involving Fireworks AI. Cursor co‑founder Aman Sanger called it an omission not to cite Kimi in the initial announcement and said the company will correct that in future disclosures.
Cursor Composer 2 Kimi K2.5 Transparency Controversy
Cursor shipped Composer 2 on March 19. Three days later a developer discovered the string kimi-k2p5-rl-0317-s515-fast in the product's API configuration, revealing that Composer 2 is built on Moonshot AI’s open-source Kimi K2.5 Mixture-of-Experts (MoE) model. The discovery sparked questions about transparency and open-source ethics. The author reports benchmark and cost comparisons: CursorBench scored Composer 2 at 61.3 with a small Terminal-Bench gap vs Claude (3.7 points); Composer 2’s input-token price is cited at $0.50/M making it ~30× cheaper than Opus 4.6 in the author’s comparison. The piece also disputes Cursor’s claim that “75% of compute was ours” and notes common developer workflows split work between Cursor (≈80%) and Claude Code (≈20%).
Cursor's Composer 2.5 Fast Outperforms Composer 2.5
Cursor released Composer 2.5 and Composer 2.5 Fast. An independent benchmark across 11 engineering skills (5 scenarios per skill, averaged over three judges) found Composer 2.5 Fast scored 92.7% with skill context versus 92.1% for Composer 2.5, completed scenarios in 59s on average versus 87s for the regular model (≈32% faster), and carries the same marginal cost under Cursor’s subscription. Both 2.5 variants outperform gpt-5.5, gpt-5.4 and Composer 2 in this evaluation. Per-skill results vary (e.g., fast wins documentation and linting; regular wins fastify, oauth, typescript), and the authors note a typescript-specific regression when using skill context.
Track Real-Time Market Signals & Shifts
Set up custom watchlists to receive automated, evidence-grounded executive digests whenever material signals or shifts occur across your tracked landscape.
