Founderland Logofounderland
the ★ top ★ 100 ★ marketers ★
SavedSearch
FoundersFounders
Fintech iconFintechClimate / Social Tech iconClimate / Social TechSaaS iconSaaSHealthtech & Biotech iconHealthtech & BiotecheCommerce iconeCommerceMedia & Entertainment iconMedia & Entertainment
Fintech iconFintechClimate / Social Tech iconClimate / Social TechSaaS iconSaaSHealthtech & Biotech iconHealthtech & BiotecheCommerce iconeCommerceMedia & Entertainment iconMedia & Entertainment
Fintech iconFintechClimate / Social Tech iconClimate / Social TechSaaS iconSaaSHealthtech & Biotech iconHealthtech & BiotecheCommerce iconeCommerceMedia & Entertainment iconMedia & Entertainment
Product Launches
Industries
Fintech iconFintechClimate / Social Tech iconClimate / Social TechSaaS iconSaaSHealthtech & Biotech iconHealthtech & BiotecheCommerce iconeCommerceMedia & Entertainment iconMedia & Entertainment
Investment News
Industries
Fintech iconFintechClimate / Social Tech iconClimate / Social TechSaaS iconSaaSHealthtech & Biotech iconHealthtech & BiotecheCommerce iconeCommerceMedia & Entertainment iconMedia & Entertainment
Research & Innovation
Industries
Fintech iconFintechClimate / Social Tech iconClimate / Social TechSaaS iconSaaSHealthtech & Biotech iconHealthtech & BiotecheCommerce iconeCommerceMedia & Entertainment iconMedia & Entertainment
FoundersFounders
Return

Recommended Articles

SaaS iconSaaSOctober 4, 2026

Subvocal launches under-chin wearable for silent computer control

Subvocal launches under-chin wearable for silent computer control
YcBrain Computer Interface+3
SaaS iconSaaSOctober 4, 2026

DoD Solution raises $2M for AI drone navigation in war zones

DoD Solution raises $2M for AI drone navigation in war zones
Defense TechDrone Tech+3
Fintech iconFintechApril 15, 2026

Colosseum's $2.75M Solana Hackathon Deploys AI Research Assistant

Colosseum's $2.75M Solana Hackathon Deploys AI Research Assistant
Blockchain InfrastructureAi Agents+2
SaaS iconSaaSApril 15, 2026

Lovable Launches One-Click Monetization for AI-Built Apps

Lovable Launches One-Click Monetization for AI-Built Apps
AiDeveloper Tools+3

Founders Mentioned

Michael undefined

cursor

saas icon
SaaS

Michael undefined

cursor

saas icon
SaaS
SaaS iconSaaS
April 15, 2026
Ai BenchmarkingDeveloper ToolsAi EthicsGenerative Ai

Cursor's Composer 2: Promising Benchmarks, Disclosure Controversy

The $29.3B AI startup launched its new coding model with impressive performance claims—then admitted it built on Moonshot AI's Kimi. Users report mixed results.

Cursor's Composer 2: Promising Benchmarks, Disclosure Controversy

The numbers looked impressive. When Cursor unveiled Composer 2 on March 19, the AI coding startup pitched it as a breakthrough in cost and performance—frontier-level results at prices that undercut established rivals by a substantial margin. Internal benchmarks showed a leap from 44.2 to 61.3. Industry tests suggested it could edge past Claude Opus 4.6, if not quite match GPT-5.4's commanding lead.

What Cursor didn't say, at least not initially, was that Composer 2 wasn't built from scratch.

Three days after launch, following pointed questions on social media, the company acknowledged what it had left out of the announcement: Composer 2 relied on Moonshot AI's Kimi K2.5 as its foundation. Cofounder Aman Sanger called the omission a "miss." A week later, Cursor published a technical report naming both Kimi and Fireworks AI, the partner that brokered the commercial arrangement. By then, though, the narrative had shifted. Forum threads and Reddit posts were describing instruction-following issues, with claims of hallucination and sluggish responses—complications that muddied the sleek launch story.

The incident raises a familiar question in enterprise AI: How much does transparency about a model's origins matter when the final product delivers? For Cursor, the answer may depend on whether Composer 2's performance stabilizes enough to justify the early hype.

The Performance Pitch

Cursor positioned Composer 2 as a rethinking of the cost-performance equation for coding assistants. On CursorBench, the company's internal evaluation, the new model scored 61.3—a significant jump from Composer 1.5's 44.2. VentureBeat cited Terminal-Bench 2.0 results that put Composer 2 at 61.7, ahead of Claude Opus 4.6's 58.0 but still trailing GPT-5.4's 75.1.

The real differentiator, though, was pricing. Composer 2 Fast—the default variant that trades some latency for higher throughput—costs $1.50 per million input tokens and $7.50 per million output tokens. The standard tier runs even cheaper: $0.50 and $2.50, respectively. Both offer a 200,000-token context window, according to VentureBeat, and integrate tightly with Cursor's environment. File operations, shell commands, browser automation—all native, all designed for the kinds of multi-step workflows that Claude and GPT handle only through external orchestration layers.

There's a catch. Unlike Anthropic or OpenAI, Cursor doesn't offer Composer 2 as a standalone API. The model lives inside Cursor's editor and an early alpha interface. It's available only to the company's existing user base, which Cursor claimed had crossed $1 billion in annualized revenue as of November 2025.

VP Lee Robinson framed it as a calculated constraint. Frontier-level coding, but locked to Cursor's ecosystem. Speed and economics, packaged as a value proposition for customers already paying for the platform.

The Disclosure Problem

Cursor's launch blog didn't mention Kimi. It highlighted benchmarks, pricing, tool integration—the usual playbook for a model rollout. The technical architecture? Absent.

On March 22, TechCrunch reported that Composer 2 had been built on top of Kimi K2.5, citing posts from Robinson and Sanger on X. Robinson acknowledged the open-source foundation and estimated that roughly one-quarter of the compute came from Kimi, with the rest from Cursor's continued pretraining and reinforcement learning work. Kimi's official account confirmed the arrangement, noting that Cursor had accessed the model through an authorized commercial partnership with Fireworks AI.

Sanger's public response was brief. A miss.

Cursor published its technical report on March 27—eight days after launch. The document explicitly credited Kimi K2.5 and thanked Fireworks. It detailed the company's approach to real-time reinforcement learning, a methodology that allows Cursor to ship updated model checkpoints to production as often as every five hours, ingesting signals from user interactions to refine behavior. A separate blog post the day before outlined this RL strategy, which Cursor had previously applied to its Tab autocomplete feature.

The acknowledgment framed the base model as a starting point rather than the core differentiator. But the timing left an impression, particularly among enterprise customers who care about training provenance and IP risk. In a market where transparency has become a selling point—Anthropic publishes model cards; OpenAI details training processes—Cursor's initial silence felt out of step.

What Composer 2 Actually Does (and How Users Say It's Working)

Digital illustration for article section "What Composer 2 Actually Does (and How Users Say It's Working)" in "Cursor's Composer 2: Promising Benchmarks, Disclosure Controversy" - A conceptual and minimalist representation of an intelligent system orchestrating multiple tasks, fe...

Composer 2 is a coding-specific model designed for long-horizon tasks inside Cursor's editor. It orchestrates file edits, executes shell commands, and handles browser automation, all through a multi-agent interface introduced in Cursor 2.0 last October. The "Fast" variant is the default setting, optimized for throughput in iterative workflows.

The technical report emphasized continued pretraining on code-specific data and real-time RL as the levers that pushed Composer 2 past its predecessor. Users access the model within their existing Cursor plans, drawing from a standalone Composer usage pool. The model competes with Claude and GPT on coding benchmarks, though Cursor's messaging emphasizes cost and Cursor-native tooling rather than raw intelligence alone.

That's the pitch. The early user reports tell a messier story.

Forum threads and Reddit posts in the weeks following launch painted a less optimistic picture. Multiple users described instruction-following issues, with Composer 2 making excessive or unexpected edits. Some reported hallucinated outputs—foreign file paths, unrelated code snippets, or content in languages they hadn't used. Cursor staff responded on March 27 to one such thread, clarifying that this wasn't a cross-user data leak but rather hallucination by a model routed in auto mode. The recommended fix: start a new session or pin a specific model like Claude or GPT. Not exactly reassuring.

Other complaints centered on performance. Users observed inconsistent cache-read behavior with Composer 2 compared to Composer 2 Fast and the older 1.5 version. User reports describe intermittent slowness and "verifying" hangs that left developers waiting. Workarounds included toggling HTTP/1.1 compatibility—a technical fix that suggests infrastructure issues, not just model behavior. Cursor's status page logged several incidents between April 1 and April 14, mostly tied to degraded performance for third-party models and elevated agent error rates, though not Composer-specific outages.

Perhaps most pointed were the comparisons to Composer 1.5. Several developers asked Cursor to bring back the older model, describing Composer 2 as unstable and prone to overwriting code without clear justification. These are anecdotes, not systematic evaluations. But they arrived frequently enough in the first two to three weeks to complicate the launch messaging.

The Bigger Competitive Stakes

Digital illustration for article section "The Bigger Competitive Stakes" in "Cursor's Composer 2: Promising Benchmarks, Disclosure Controversy" - A clean and minimalist conceptual artwork representing a high-stakes competition in a crowded market...

Cursor operates in a crowded field. Anthropic publicized strong Claude Code revenue run-rate growth in its Series G coverage in early 2025, setting a baseline for what agentic coding tools can monetize. GitHub Copilot, Replit, and a dozen smaller players are all racing to prove they can replace or augment human developers at scale. Cursor's November 2025 Series D—$2.3 billion at a $29.3 billion valuation—positioned the company as a leader, with NVIDIA and Google joining as strategic investors. NVIDIA is also cited as an enterprise customer, a detail that matters in a market where credibility with Fortune 500 buyers counts.

The revenue trajectory supports the valuation. As of June 2025, Cursor claimed more than $500 million in annualized revenue and usage by over half the Fortune 500, including NVIDIA, Uber, and Adobe. By November, that ARR figure had crossed $1 billion, with a team of over 300. Not bad for a startup that didn't exist in 2023.

Composer 2 fits into this arc as Cursor's second in-house model, following the original Composer's October 2025 debut. The company acquired Supermaven in November 2024 to strengthen code completion, and the Composer line represents its push into multi-agent orchestration. Real-time RL—the ability to ship improved checkpoints every few hours based on live user feedback—is the technical differentiator Cursor emphasizes most. It's also what the company argues sets Composer 2 apart from the Kimi base model.

The disclosure complication muddies that narrative. Building on an external base isn't unusual in AI. Even OpenAI has acknowledged using third-party data and tooling in its training pipelines. Cursor emphasizes its RL methodology and continued pretraining as substantial work that differentiates Composer 2. But the omission at launch raised questions about transparency in a market where provenance and training practices matter to customers, especially enterprise buyers evaluating compliance and IP risk.

Cursor has addressed the gap, if belatedly. The technical report is public, and the company has been vocal about its RL methodology. Whether that's enough to restore trust depends on whether Composer 2's performance, once bugs stabilize, justifies the benchmarks.

Early user feedback suggests that's still an open question. The model may deliver on cost and speed. But in a market where developers have choices—and where reputation compounds quickly, both good and bad—Cursor's handling of the launch left an opening. Whether competitors exploit it depends on how fast the company can turn anecdotes into artifacts: stable releases, consistent performance, and the kind of reliability that justifies a $29.3 billion valuation.

For now, Cursor is shipping updates. The real test is whether they arrive fast enough.

More stories

  • Subvocal launches under-chin wearable for silent computer control
  • DoD Solution raises $2M for AI drone navigation in war zones
  • Colosseum's $2.75M Solana Hackathon Deploys AI Research Assistant
  • Lovable Launches One-Click Monetization for AI-Built Apps
  • Wavelet Medical Raises $7M Seed for AI-Powered Fetal Brain Monitoring
  • Spiral Therapeutics Lands $27M Series B for Inner Ear Treatments
fintech icon
climate-social-tech icon
saas icon
healthtech-biotech icon
ecommerce icon
media-entertainment icon
Loading...

About

Dreamwell AIContact UsOur Story

Articles

Product LaunchesInvestment NewsResearch & Innovation

founderland

We Use Cookies

We baked up some cookies – the digital kind. They help Draper run like a well-oiled mid-century machine. Some are essential to the experience, others help us tailor things to your taste. We promise, no crumbs on your blazer. Take a moment to choose what works for you.