Safi Shamsi's pitch is simple enough: Stop feeding your entire codebase into ChatGPT and hoping for the best.
On July 1, his London-based startup Graphify Labs—barely two people, fresh from Y Combinator's Summer 2026 batch—released an open-source engine that maps software repositories into knowledge graphs. The idea is to give AI coding assistants something closer to a blueprint than a haystack. Engineers at Rootly are already running it in production, according to the company, while Geotab and Tweddle Group are listed as users in the YC directory. And by one measure at least, the tool has found an audience: PyPI, the Python package index, has logged roughly 3.6 million downloads.
Whether that translates to actual staying power is another question entirely.
The Context Problem
The release arrives at a particular moment in the AI tooling cycle. Development teams are racing to adopt assistants like GitHub Copilot, Cursor, and Claude Code. Most of these tools work by searching through files—either with old-fashioned grep or newer vector embeddings—and then stuffing whatever they find into the model's context window. The approach works, sort of. But it misses the relationships that matter: which function calls which, how classes inherit from one another, where imports cascade through a project.
Graphify's answer is deterministic rather than probabilistic. Run /graphify . in a repository and the tool parses the abstract syntax tree using Tree-sitter grammars. No cloud inference, no telemetry. Everything happens on the machine. The output is a typed graph—functions, classes, files become nodes; calls, imports, definitions become edges. The assistant queries the graph instead of fumbling through raw text.
The tool supports 36 languages out of the box. Python, JavaScript, Rust, Go, and on down the list. For non-code files (PDFs, images, audio transcripts), there's an optional semantic pass that sends snippets to a model of your choosing—Claude, OpenAI, Gemini, DeepSeek. The graph exports to Neo4j, FalkorDB, GraphML, or Obsidian, useful for teams with existing workflows they'd rather not abandon.
Integration happens through two paths. Some assistants, like Claude Code or Cursor, use hook-based interception that injects graph context before the model sees the prompt. Others connect via the Model Context Protocol, an emerging standard that treats graph operations—query_graph, get_node, shortest_path—as callable functions. The MCP server runs over stdio or HTTP, which means it works in IDEs or remote setups without much fuss.
Verification at the Merge Gate

The enterprise layer, currently in early access, adds a feature that sounds almost academic: differential verification at pull request time. When a developer submits a change, Graphify uses theorem provers—Z3, CBMC, JBMC, CrossHair—to check whether each modified function behaves identically to the original or has changed semantics. Functions that pass equivalence checks get waved through. Those that don't get flagged for human review.
The verification is bounded, which is to say it proves properties within defined limits: loop iterations, recursion depth, buffer sizes. When proof paths exceed those constraints, the system abstains rather than guessing. Java and Python verification relies on symbolic execution with similar boundaries. All of it happens on-premise. No code leaves the building, a detail that matters to compliance-heavy industries.
Other enterprise features include graph-aware code review (which surfaces downstream impact for proposed changes) and an engineering digest that summarizes structural shifts over time. Pricing isn't public. The enterprise tier is being rolled out through a waitlist, the usual startup playbook.
Stars, Downloads, and the Credibility Gap
Graphify's GitHub repository crossed 96,432 stars on launch day, according to a July 1 company blog post. By late July, that figure had climbed above 96,700. The velocity raised eyebrows. Stridenote, a developer analysis site, published a code inspection on July 11 suggesting readers "triangulate impact via forks, issues, releases, community tutorials rather than stars alone."
Fair advice, perhaps. GitHub stars are easily gamed and rarely correlate cleanly with actual usage.
The PyPI package—confusingly named graphifyy with a double-y, a quirk Stridenote flagged as "typosquat-shaped"—shows 3.6 million all-time downloads and 1.4 million in the last 30 days, according to pepy.tech data from late July. Those figures include CI traffic and mirror pulls, so they're directional at best.
Rootly, an incident management platform, published a blog post in April about a Graphify plugin for turning incident data into a knowledge graph. The integration appears on Graphify's customer page. Geotab and Tweddle Group are listed in the YC directory as production users, though neither company has published a case study or public confirmation as of late July.
Community sentiment is split. Reddit threads from late June and July show some developers reporting token savings and faster assistant queries. Others flag compatibility issues with specific Claude Code versions or complain that the assistant ignores the graph tools in practice. A May 28 review by RoboRhythms claimed savings ranging from 6.8× to 49× typical, though the author noted a hook regression in Claude Code v2.1.117 that broke automation for some users.
In other words, the tools appear to work for some teams under some conditions. Not exactly a ringing endorsement, but not damning either.
An Increasingly Crowded Space

Graphify is launching into a field that's filling up fast. Sourcegraph's Cody has been using an internal code graph for context since 2023. Newer open-source alternatives like code-review-graph (MCP-native, SQLite-based) and GitNexus (browser-based, client-side) are releasing updates weekly. An academic paper from March described Codebase-Memory, another Tree-sitter-based knowledge graph built on MCP. Local parsing plus graph traversal is becoming table stakes, not a differentiator.
What Graphify is betting on, at least publicly, is breadth: 36 language grammars, 17 assistant integrations, and that enterprise verification layer. Whether that's enough depends on how well the tools actually perform in production environments at scale—a question most buyers don't have a clean answer to yet.
The model isn't new. Developers have been building code intelligence tools for decades, from ctags to language servers to modern IDE analyzers. What's changed is the urgency. AI assistants need better maps, and companies are willing to experiment with anything that promises to cut token costs or improve accuracy.
What Comes Next

Shamsi, the founder and CEO, previously worked as an AI engineer at Valent Projects and holds a master's in data science from the University of Birmingham. The company's about page describes him as the author of "The Memory Layer," a thesis on knowledge graphs in retrieval-augmented generation, though that work hasn't been independently verified here. The team is hiring, but no formal job listings were posted as of late July.
The core product is available now under an Apache 2.0 license. (The repo also contains a separate MIT license file for certain components, though the precise scope isn't clarified in a single document.) Enterprise features are rolling out through early access. The YC S26 Demo Day is scheduled for September 10, which may bring additional product announcements or funding news.
For engineering leaders evaluating code intelligence tools, Graphify represents a bet on structure over similarity, local-first over cloud-native, graph traversal over vector search. The company's challenge is converting GitHub stars and PyPI downloads into sustained enterprise adoption—a gap that typically requires customer proof points, not just open-source momentum.
The tools are live. The verdict, as they say, is still being written.
