The pull requests keep arriving. Entire features, written by AI coding agents in minutes instead of days, land in review queues stripped of their origin stories. No comments explaining the architectural choices. No commit messages capturing the back-and-forth that led to this particular implementation. Just diffs—raw, context-free diffs—waiting for human engineers to decode.
On February 10, 2025, a startup called EntireHQ surfaced with both a diagnosis and a proposed cure. The company emerged from stealth mode with $60 million in seed funding (a remarkable haul for an opening round) and a single product: Checkpoints, an open-source tool that treats AI coding sessions as permanent Git artifacts, complete with the reasoning that produced them.
The premise sounds almost obvious once you hear it. If AI agents are generating production code, shouldn't their thought process travel with it?
Versioning the Conversation
Checkpoints operates as a command-line interface that captures everything an AI coding agent does during a work session—the prompts fed to it, the full conversation transcript, which tools it invoked, which files it touched, how many tokens it consumed. Then it links all that metadata directly to Git commits as versioned data.
Thomas Dohmke, EntireHQ's founder, framed the tool in his LinkedIn launch post as "a new primitive that automatically captures agent context as first-class, versioned data in Git." Perhaps a touch of startup hyperbole there, but the underlying mechanic is more straightforward than the buzzwords suggest.
Here's how it works: Entire installs hooks into Git and your AI agent. As the agent operates, the system saves sessions and checkpoints to a separate branch—entire/checkpoints/v1—keeping your main code history uncluttered while tying commit SHAs to the captured context. At launch, Claude Code from Anthropic and Google's Gemini CLI (in preview) are the supported agents.
The CLI offers two capture modes, each with trade-offs. The default, manual-commit mode, creates permanent checkpoints when you run git commit. It quietly appends an "Entire-Checkpoint: <id>" trailer to your commit message—no extra commits required on your working branch. Later, when reviewers pull up that commit, they can click through to see what the AI agent was thinking.
Auto-commit mode takes a more aggressive approach, automatically creating a commit after each agent response that modifies files. These commits carry the prefix "[entire] …" which, as the documentation candidly notes, can clutter your branch history something fierce. The company recommends using auto-commit only on feature branches and squashing before you merge to main. Fair advice, that.
Each checkpoint belongs to a broader session, which stores the complete conversation transcript, tool calls with their arguments, token usage, and even per-line attribution for who (or what) wrote which code. The system can handle nested sessions when one coding agent spawns sub-agents, capturing what Entire calls "the entire tree of reasoning." (The company may have been established solely to make that pun work.)
The Review Problem Gets Real

Why does any of this matter? Because engineering teams have discovered a new bottleneck in their development workflows, one that wasn't on anyone's radar two years ago.
AI coding agents can generate features faster than teams can review them. According to TechCrunch's coverage, Checkpoints aims to "pair every bit of software the agent submits… with the context that created it, including prompts and transcripts." GeekWire described it as recording "the reasoning and instructions behind AI-generated code and saving that information together with the code itself."
The problem isn't just velocity. It's governance. When AI-generated code ships to production and something breaks six months later, how do you audit what happened? How do you understand the intent behind architectural decisions when no human was in the room?
Entire's web dashboard addresses part of this. It displays checkpoints by branch, showing checkpoint IDs, timestamps, owners, which agent did the work, files changed, and token counts. Pull request commit messages include checkpoint IDs with deep links into the Entire platform. Reviewers get side-by-side diffs, code churn metrics, a sessions panel with full transcripts, and metadata linking back to the original commit.
The CLI includes rewind functionality—commands that let you roll back to any checkpoint, either interactively or programmatically. The tool warns that rewinding discards post-checkpoint changes unless you've saved them elsewhere, a sensible guardrail. There's also an "entire explain" command that can show or generate summaries (requiring Claude CLI) and display parsed transcripts for any checkpoint or commit.
Optional AI-generated summaries can run at commit time, configurable through strategy options. Privacy controls include a telemetry flag and the ability to disable anonymous usage analytics. Because nobody wants their internal codebase conversations leaked into someone else's training data.
Git-Native Versus Session-Scoped
Other AI coding tools already use the term "checkpoints." Cursor IDE has them. So do Claude Code, Gemini CLI, Salesforce Agentforce, Kilo Code, and Conductor. But those implementations typically function as local, session-scoped snapshots—safety nets for the current work session, not permanent fixtures in your repository's history or pull request workflows.
Entire's Checkpoints, by contrast, are agent-agnostic and commit-addressable. The metadata lives in Git itself, not some isolated session store that vanishes when you close your IDE. That architectural choice matters for organizations with compliance requirements or teams that need to audit AI-generated code months or years after deployment.
The distinction between session-scoped and Git-native may sound technical, and it is. But it's also strategic. Session-scoped checkpoints help individual developers recover from mistakes. Git-native checkpoints create a paper trail.
On Hacker News, where developers tend to be both enthusiastic and brutally honest about new tools, the Checkpoints announcement sparked discussions about prior attempts to store AI prompts and intent alongside code. The thread validated the concept's importance—perhaps reflecting a shared frustration with the current state of AI-assisted development that nobody quite knew how to articulate yet.
Dohmke positioned Checkpoints as the "first crack" at what he envisions as a universal semantic reasoning layer, something that will eventually become shared memory for agent collaboration. SiliconANGLE reported the system logs prompts and token usage to help teams reuse troubleshooting prompts and avoid making the same mistakes twice with their AI agents.
Ambitious language, certainly. Whether the reality lives up to that vision remains to be seen.
Shipping Before Perfect

Entire hasn't built its own coding agent, a choice that Axios noted in their coverage. Instead, the company integrates with existing agents—a pragmatic move in a fragmented AI tooling landscape where new models and interfaces appear monthly.
The web UI and the promised semantic layer remain early-stage efforts. The February 10 launch focused on shipping the CLI and proving the Git-native capture mechanic actually works. The Gemini CLI integration carries labels like "preview" and "work in progress" in the documentation. Claude Code is the primary first-class integration, which tracks with Anthropic's strong position in developer tools.
For teams interested in trying Checkpoints, the barrier to entry is low. The tool is open source under an MIT license, available in the entireio/cli repository on GitHub. Installation follows familiar patterns: brew tap entireio/tap && brew install entireio/tap/entire on macOS or Linux, with Windows support via WSL. Go users can opt for go install github.com/entireio/cli/cmd/entire@latest.
Prerequisites are minimal: Git and an AI coding assistant (Claude Code or Gemini CLI). The CLI commands read like standard developer fare: entire enable, entire rewind, entire resume, entire explain, entire status, entire doctor, entire reset. Documentation lives at docs.entire.io, covering strategies, commands, integrations, and the web UI. Recent GitHub release activity shows frequent updates—changelogs reference checkpoint metadata improvements, Gemini hooks, summaries, and commit trailers.
The core idea is out there now, open source, usable today. Whether Checkpoints becomes the standard way to version AI reasoning or just another well-intentioned experiment in an increasingly crowded field depends entirely on adoption patterns over the coming months.
For engineering teams already drowning in agent-generated pull requests, the promise of reviewable context might be compelling enough. Or it might be one more tool to learn, one more configuration to manage, one more layer in an already complex development stack.
Time, as they say, will tell. But at least someone's trying to solve the right problem.
