Picture an artificial intelligence agent tasked with finding land for a new data center. Simple enough, right? The reality is more Kafkaesque than most technologists care to admit.
That hypothetical agent needs flood maps from FEMA, soil surveys buried in USDA databases, transmission line distances scattered across utility portals, and zoning ordinances from—brace yourself—one of roughly 33,000 local jurisdictions across America. Each dataset speaks a different digital language. Each updates on its own inscrutable schedule. And for an AI designed to work autonomously, this isn't merely frustrating. It's paralyzing.
Enter Mireye, a two-person outfit fresh from Y Combinator's Summer 2026 batch. The San Francisco startup has assembled what it calls "infrastructure for Physical World AI Agents"—aggregating more than 300 geospatial fields from 85 authoritative sources into a single API. Early customers claim they're sourcing off-market land for data centers "100x faster." Whether that figure holds up under scrutiny remains an open question, but it hints at the scale of inefficiency Mireye is targeting.
When Silicon Valley Collides With Dirt
Jensen Huang, NVIDIA's perpetually leather-jacketed CEO, has been evangelizing "Physical AI" in keynotes from 2024 to 2026—agents that don't just generate clever prose but interact with buildings, power grids, supply chains. By 2026, that vision is meeting messy reality.
Microsoft's Work Trend Index, published in April, documented broad agent adoption across industries through surveys and usage data. Forrester's June report noted companies "chasing" agentic AI even as governance incidents accumulate. The infrastructure supporting this shift? Lagging badly.
A March research paper on the "Internet of Physical AI Agents" outlined the need for policy-governed runtimes and interoperability frameworks—academic speak for "we're building agents on duct tape and prayers." Another study from the same period showed autonomous agents slashing affordable housing site selection in New York from 18 months to 72 hours. Impressive, until you learn researchers first spent weeks wrangling data on regulatory constraints, zoning, and demographics from dozens of disconnected sources.
The data layer, in other words, is a disaster.
Data Centers as Proving Ground
Nowhere is this fragmentation more acute than in data center siting, an industry suddenly facing existential constraints.
Gartner forecast in June that data center electricity consumption would hit 565 terawatt-hours in 2026, up from 447 the year before. AI-optimized servers now claim 31% of that power draw. Newmark reported a US pipeline of roughly 160 gigawatts under construction or announced. Demand, in short, is ferocious.
But the grid can't keep pace. In August, Texas imposed a moratorium on new data center interconnections after requests hit 474 GW—roughly five times the state's typical peak demand. FERC launched targeted actions in June to reform how regional transmission organizations handle large load interconnections, a direct response to the backlog hyperscale AI infrastructure created.
Finding viable sites now means orchestrating data on transmission proximity, substation capacity, interconnection queues, flood zones, wetlands, zoning ordinances, and a dozen other layers. Each maintained by different agencies. Each operating on different timelines. Berkeley Lab's "Queued Up: 2026 Edition" analyzed interconnection queues representing the vast majority of US generation capacity, yet accessing that data alongside land-use constraints still requires navigating a labyrinth of fragmented portals.
It's enough to make even seasoned developers long for simpler times.
The Aggregation Layer

Mireye's answer is a REST API and an MCP (Model Context Protocol) server connecting Claude, ChatGPT, Gemini, and other agents to what the founders describe as "cited, provenance-rich physical-world data." The catalog spans 300-plus fields covering land, location, people, rules, and risks.
Here's where things get interesting: each field returns not just a value but full source attribution. The originating agency. A URL. A timestamp of when it was fetched. A confidence score. For regulated use cases—insurance underwriting, commercial lending, energy siting—that paper trail matters. A lot.
The API exposes endpoints for LLM-backed Q&A with per-citation provenance, batch lookups, proximity analysis over infrastructure like airports and substations, and field requests that queue new data builds when information is missing. That on-demand indexing feature resembles a marketplace flywheel, absent from most static geospatial catalogs.
Pricing follows a credit model: 5,000 credits free monthly, paid tiers starting at $19, additional credits at $1 per 1,000. Coverage spans the US, with proximity queries extending into Canada. Sources include the usual federal suspects—USGS, FEMA, NOAA, USDA, EPA, EIA, FCC, Census, NREL—alongside commercial providers like Overture Maps and Regrid.
A demo on the company's site shows an elevation lookup sourced to USGS 3DEP Cloud Optimized GeoTIFF, fetched July 28, 2026. Clean. Traceable. Exactly what an agent—or the regulator auditing its decisions—would need.
The Fragmentation Beneath

Understanding why Mireye exists requires appreciating just how byzantine American governance is.
Roughly 3,144 counties. Over 33,000 zoning jurisdictions that the National Zoning Atlas initiative aims to digitize. Around 2,000 public power utilities. Another 900 electric cooperatives covering 56% of the nation's land. Each layer maintains its own datasets, often in legacy formats or behind APIs designed for human analysts rather than autonomous agents.
Federal datasets are migrating toward cloud-native formats—COG for raster data, GeoParquet for vector layers—but unevenly. USGS serves 3DEP elevation as COG. NASS released the 2025 Cropland Data Layer with Hawaii updates as recently as June 12, 2026. NRCS refreshed SSURGO soil surveys in November 2024 with SQLite downloads. FEMA's National Flood Hazard Layer comes via ArcGIS FeatureServer, though data currency varies by community. The FCC's National Broadband Map fabric, available via APIs and bulk downloads, was still clarifying data handling rules in June through DA 26-630.
Then there's licensing. Overture Maps uses CDLA-Permissive-2.0 where possible, but some themes remain under OSM's ODbL share-alike license due to OpenStreetMap inputs. Derivative databases must navigate these mixed-license mosaics—something agents need to respect but aren't naturally equipped to parse.
It's a mess. An understandable mess, given the federal structure of American governance, but a mess nonetheless.
The MCP Bet
Anthropic introduced the Model Context Protocol in November 2024, then donated it to the Agentic AI Foundation under the Linux Foundation umbrella in December 2025. By mid-2026, MCP servers are proliferating as gateways between domain APIs and LLM agents.
Mireye's MCP server, listed on Glama and distributed via PyPI as mireye-earth-mcp, proxies the HTTP API for agent frameworks. Research published between May and July 2026 catalogs both rapid ecosystem growth and emerging security concerns. One study flagged vulnerabilities in MCP deployments accessing critical infrastructure data. Another outlined design patterns for human-agent collaboration architectures.
The protocol enables agents to deterministically call domain APIs, sure. But it's also exposing gaps in server security posture that enterprises are only beginning to grapple with.
For Mireye, MCP isn't just technical plumbing—it's a positioning bet. Generic geospatial platforms like CARTO and Foursquare Studio offer analytics and visualization, but they're optimized for human workflows. Mireye's provenance envelopes—source, timestamp, confidence per field—target regulated use cases where an agent's decision trail must be auditable. Get the data lineage wrong in insurance underwriting or commercial lending, and liability follows.
Crowded Territory
Mireye isn't alone in spotting the opportunity, naturally.
National Flood Data offers API-first risk data for underwriting automation. Zoneomics provides a zoning API for the US and Canada with last_updated fields per zone. LightBox's SmartParcels and SmartFabric bundle parcel data with optional zoning add-ons. Regrid maintains a nationwide parcel API with pricing via their store. ATTOM Data serves property details, comparables, tax, and lien data. Mapped positions itself as a "world model for Physical AI" across facilities and sensors, complete with a data center industry page.
What differentiates Mireye appears to be breadth across domains—not just parcels or flood risk, but the full constraint stack—plus agent-native design through the MCP server and citation metadata, and that on-demand indexing queue. Whether any of this proves defensible depends on execution speed and how quickly incumbents adapt their APIs for agent consumption.
Incumbents, after all, have a nasty habit of waking up.
Regulatory Winds

The EU AI Act entered staged enforcement in 2026. Obligations for foundation models kicked in August 2; high-risk Annex III systems face compliance in December 2027; embedded systems under Annex I get pushed to August 2028. Provenance and audit requirements favor systems that can prove data lineage, vintage, and policy-bounded tool use—precisely the metadata Mireye wraps around each field.
Stateside, FERC's Orders 2023 and 2023-A, finalized March 21, 2024, reformed interconnection procedures. The June 2026 large-load docket added procedural clarity for generators and AI data centers. CEQ's Phase 2 NEPA final rule from May 1, 2024, aimed to streamline environmental review, creating a regulatory environment that rewards data-rich, defensible siting workflows.
The geospatial standards landscape is consolidating, too. The OGC approved API – Features Part 3 (Filtering) as an official standard, nudging the ecosystem toward developer-friendly OpenAPI patterns. COG and GeoParquet now carry endorsements in 2026 Federal Geographic Data Committee and National Geospatial Advisory Committee recommendations. Overture Maps' June 17 release continued the shift to GeoParquet, with federal portals following.
Standards battles are dull until they're not. And this one's tilting toward machine-actionable formats.
Three Forces to Watch
Whether unified data layers for physical AI become critical infrastructure or remain niche tools hinges on three dynamics.
First: continued migration of federal and open datasets to cloud-native formats with machine-actionable metadata. The faster COG and GeoParquet adoption spreads, the less manual wrangling startups like Mireye must do—and the more defensible provenance-tracking becomes as a value layer.
Second: whether data center and renewable siting pressures validate the "100x faster" claims. The Texas moratorium and FERC's June actions signal that grid interconnections, not land availability, are the binding constraint. If Mireye's proximity queries over substations and interconnection queues genuinely shorten development timelines, traction will follow. If they surface sites that still take years to interconnect, the value proposition narrows considerably.
Third: MCP ecosystem maturity. As enterprises expose more domain APIs as MCP servers—science data, broadband infrastructure, GIS layers—agent-grade connectors will either commoditize or fragment along vertical lines. Mireye's on-demand indexing could become a moat if field requests from one customer benefit all others, creating network effects. Or it could be table stakes if every competitor builds a similar queue.
The Market Underneath
The geospatial analytics market is sizable, though market research firms can't quite agree on the numbers. Grand View Research pegged it at $102.7 billion in 2025 and $117.2 billion in 2026, projecting $234 billion by 2033. Fortune Business Insights forecast $117.30 billion in 2026 growing to $309.84 billion by 2034.
Those figures lump together traditional GIS, remote sensing, and location intelligence—not just agent-facing APIs. Still, they hint at underlying demand for spatial reasoning as infrastructure development accelerates and AI agents graduate from chatbots to decision-makers.
For now, Mireye is a team of two in San Francisco with a catalog of 300 fields and a thesis: physical AI agents need more than raw data. They need cited, timestamped, confidence-scored facts they can defend to regulators, lenders, and insurers.
Whether that's enough to build a platform business or just a feature some larger player eventually acquires? Ask again in a year, perhaps when the first wave of agent-driven siting decisions either close or collapse under regulatory scrutiny. The infrastructure for Physical AI is being assembled in real time, one API endpoint at a time. Mireye is betting it can lay the foundation before the opportunity ossifies into commodity plumbing.
Smart bet or premature optimization? The physical world, as always, will have the final say.
