A license plate glides under the overhang at a car wash entrance. The camera mounted above doesn't just record the moment—it recognizes the plate, checks it against a loyalty database, applies the member's discount in the point-of-sale system, logs the visit, updates their profile, makes the transaction searchable. All of this happens in the seconds before soap hits windshield. No cashier. No app. No human in the loop at all.
This, more or less, is the future being sold right now to hundreds of millions of existing cameras: a shift from passive observer to autonomous actor.
OpenVector, a startup emerging from computer vision research at Harvard, is among those racing to define what that future looks like. Co-founders Andrey Gizdov—a Fulbright scholar who walked away from his PhD—and Vishal Urlam describe their pitch simply: "an AI that lives in your camera to take action." Connect any existing camera to business systems through plain-English commands. No hardware replacement required. On their website, a live demonstration lets visitors type instructions like "detect tailgating" or "identify wash club members" and watch automated workflows trigger across operational systems—DRB SiteWatch, Twilio, webhooks, the usual suspects.
The underlying premise isn't particularly exotic anymore. What's changed is that the technology actually works now, the hardware costs have collapsed, and the market—perhaps more than the founders expected—seems ready.
An Installed Base Waiting for Intelligence
The raw numbers sketch the landscape plainly enough. Outside China, the installed base of cameras had reached 562 million by late 2025, according to data from Novaira Insights cited by Axis Communications. Projections suggest that figure climbs to 736 million by 2029. Most of those cameras already sit in place—offices, warehouses, restaurants, gyms, retail floors—streaming RTSP feeds that someone occasionally reviews, if they review them at all.
The video analytics market itself has been scaling fast. Grand View Research estimated it at $12.7 billion in 2024, with projections climbing toward $18.8 billion in 2026 and $37.8 billion by 2030, implying a compound annual growth rate hovering near 19.5%. Fortune Business Insights offered slightly different figures: $14.81 billion for the current period, expanding to $65.08 billion by 2034. The broader computer vision market, meanwhile, was pegged at around $23.6 billion in 2025 and could hit $101.5 billion by 2033, per Grand View's longer-range forecast.
But raw growth in camera shipments doesn't tell the full story. What's driving the expansion isn't just more cameras. It's what those cameras can now do.
Axis, a leading manufacturer, noted in its 2026 "Perspectives" report that customers are increasingly repositioning cameras from pure security—watching for theft, monitoring perimeters—to business intelligence and operational efficiency. Gartner data from around the same period suggested that more than 55% of deep neural network analysis was happening at the point of capture, meaning on or near the camera itself. The edge, in other words, is becoming the brain.
The Academic Roots
OpenVector's technical foundation rests on a June 2025 paper published at CVPR, the premier computer vision conference. Titled "Seeing More with Less: Human-like Representations in Vision Models," the work—co-authored by Gizdov alongside researchers from Harvard and the Weizmann Institute—earned a Spotlight designation, reserved for the top 3% of submissions.
The core insight borrows from biology. Human vision doesn't process every pixel equally. Our eyes use foveated attention, concentrating detail in a small central region while leaving the periphery blurred. The paper by Gizdov et al. demonstrated that a vision-language model could retain up to 80% of its full-frame performance while analyzing just 3% of the pixels, guided by an attention-aware "sub-sampler" that decides which regions actually matter. On standard benchmarks—GQA, SEED-Bench, VQAv2—the approach delivered accuracy gains of 2% to 2.7% compared to uniform sampling under tight pixel budgets.
The practical claim: 45 kilobytes per 720p frame at that 3% pixel budget, representing a 5 to 10 times reduction in bandwidth consumption. For real-time video inference over commodity networks—especially edge deployments where bandwidth is constrained or expensive—that difference can separate viable from impossible.
OpenVector's research page explicitly frames the technique as solving the "pixel bandwidth/latency bottleneck" for multimodal perception at the edge. Gizdov chose to commercialize rather than finish his doctorate. His LinkedIn posts from the period reference the CVPR award and the decision to "develop OpenVector" instead of continuing a funded PhD track. It's a bet that peer-reviewed efficiency gains can translate into defensible product advantage in what is, unmistakably, a crowded and rapidly filling market.
The Product, Simplified

The product itself is easier to describe than the underlying mathematics. OpenVector's "how it works" page outlines a four-step flow: write a plain-English rule, the system converts it into a computer vision task, the AI monitors camera feeds for that condition, and when triggered it executes actions across connected systems.
The car wash example is specific, not hypothetical. "When a car's plate matches a wash-club member at the Entry Lane, OpenVector applies their loyalty discount in DRB SiteWatch, then logs the visit, updates the member profile, and makes it searchable." DRB's SiteWatch is a widely deployed point-of-sale system in tunnel car washes, with documented AI license-plate recognition and loyalty integration stretching back to at least 2024. OpenVector positions itself as the orchestration layer—the software that ties camera perception to those downstream actions.
The company is accepting pre-orders for an edge hardware appliance called the OpenVector Station, built around Nvidia's Jetson AGX Thor. Thor, a Blackwell-based chip that became generally available in late 2025, delivers 2,070 FP4 TOPS—roughly 7.5 times the AI compute of its predecessor, Orin—and is designed to run multiple generative models concurrently at the edge. Pricing sits at $449 per month on subscription or $6,499 up front, with on-site setup offered. The company is targeting Fall 2026 delivery in California and collecting $100 refundable deposits.
Whether OpenVector has customers beyond those pre-orders isn't disclosed on the public site. The team page lists the two co-founders and affiliations with Google, MIT, and the Weizmann Institute. There's no mention of funding, investors, or revenue. It's early, in other words—perhaps very early.
A Suddenly Crowded Field
OpenVector isn't operating in a vacuum. The day-to-day reality is that multiple companies are now pitching some variation on "turn your cameras into AI agents," and several are substantially ahead in funding and customer count.
Spot AI announced its "Video AI Agents" platform on July 27, 2026—just days ago. The company reports nearing $100 million in total funding and claims 1,000 customers across 17 industries, framing its offering as software that "turns any camera into a Video AI Agent" capable of taking "action like a person" across business systems. Spot's messaging emphasizes deployment speed—it cites a YMCA rollout completed in two weeks—and the shift "from hardware-centric surveillance to software-centric video intelligence."
Ambient.ai, which markets itself as an "agentic physical security" vendor, reported doubling new annual recurring revenue in a February 18, 2026 update. The company promotes edge-optimized vision-language models with temporal reasoning, emphasizing prevention over reactive detection. Like OpenVector, Ambient works with existing cameras and positions itself as an intelligence layer, not a hardware replacement.
Established players are adapting, too. Genetec, a long-time leader in video management software according to Omdia and Novaira rankings, began rolling out AI-assisted investigation features in its Security Center SaaS platform starting February 2026. Axis Communications launched a slate of edge-AI cameras at ISC West 2026—license plate recognition upgrades, thermal units, 360-degree models—with availability from June through August. These cameras embed analytics directly, bypassing the need for a separate appliance in many scenarios.
Rhombus Systems introduced "Custom Events" powered by large language models in May and June 2026, allowing users to create detection scenarios with natural-language prompts. Eagle Eye Networks promotes natural-language video search across its cloud VMS. Verkada announced "AI-Powered Deterrence" in February 2026, using large vision, audio, and language models to analyze scenes and escalate context-aware responses automatically.
Then there are the smaller entrants, multiplying rapidly. Sentinel (joinsentinel.com) advertises itself as turning "any camera into an AI security guard you can talk to," claiming more than 20 deployments. MAKRR positions as no-code computer vision for existing CCTV feeds. IronYun's VAIDIO solution, detailed in a March 2026 HPE brief, similarly targets IP cameras with broad analytics. Nsight Sentinel markets an on-premises AI video processor with sub-second local processing over ONVIF and RTSP streams.
The sheer density of competitors points to a market shift, not startup hype. PTZOptics partnered with Moondream on a "Visual Reasoning" initiative in February 2026, pairing robotic cameras with lightweight open-source vision-language models. Even sensor consolidation is underway: lidar maker Ouster acquired vision company StereoLabs in February 2026, citing StereoLabs' edge AI competence as a draw.
Hardware Constraints, Getting Lighter

The hardware to enable all of this is getting cheaper and more capable, though real constraints remain. Nvidia's Jetson AGX Thor, the chip powering OpenVector's edge box, represents a material step forward. Academic papers published from March through July 2026—covering Vision-Language-Action model bottlenecks, edge inference frameworks, quantization for small VLMs—demonstrate that on-device multimodal reasoning is viable on Jetson-class hardware. Performance, however, depends heavily on model size, frame rate, and task complexity.
Nvidia's own developer guidance, published in February and July 2026, recommends moving to AGX Thor when deployments require strict real-time guarantees, large models, or multiple concurrent inference streams. Translation: Orin and older chips can handle many use cases. Thor opens the door to richer, faster, or more parallel workloads.
Bandwidth is the silent bottleneck that no one outside the industry thinks about. Video streams are heavy. A 720p feed at 30 frames per second, encoded with H.264, typically runs 1 to 4 megabits per second. Multiply that by dozens or hundreds of cameras and network infrastructure becomes a limiting factor, especially in retail, warehouses, or distributed facilities. OpenVector's foveated sampling—45 kilobytes per frame, claiming a 5 to 10 times bandwidth reduction—addresses this directly. Whether those lab results hold up in messy real-world deployments remains an open question.
Gartner weighed in with a June 2026 brief arguing that cloud-only architectures hinder agentic AI performance and recommending hybrid edge processing instead. The firm projected $37.5 billion in AI-optimized infrastructure spending for 2026, with 55% allocated to inference. Much of that inference is moving closer to where data originates—cameras, sensors, devices—to reduce latency and egress costs.
Privacy and compliance aren't afterthoughts; they're architectural constraints baked into every design decision. On-device anonymization (companies like brighter AI, acquired by Milestone in April 2025, have been vocal advocates) and split-inference designs that keep raw video local are emerging as standard patterns. Academic work from 2025 and 2026 on privacy-preserving video analytics reflects growing regulatory and customer pressure to get this right.
Regulatory Currents, Shifting Fast
Policy is shaping who can deploy what, and where. In the United States, the FCC's Secure Equipment Act and a 2022 order (FCC 22-84) prohibit authorization, marketing, and import of surveillance equipment from specific Chinese manufacturers, including Hikvision and Dahua. The National Defense Authorization Act's Section 889, in force since 2019 and reinforced by GSA guidance as recently as April 2025, bans federal procurement of covered equipment.
The practical effect: enterprises and government agencies with existing Chinese-made cameras face rip-and-replace cycles, creating an opening for software-first solutions that work with any camera brand. OpenVector and its competitors benefit indirectly—they're hardware-agnostic overlays, not locked to specific OEMs.
The FTC's December 2023 settlement with Rite Aid over facial recognition—a five-year ban on the technology for surveillance purposes—sets a cautionary precedent for biometric deployments. The case underscores compliance obligations around algorithmic accuracy, error handling, and transparency, requirements that extend well into 2028 and beyond.
Europe's AI Act entered force on August 1, 2024, with staged implementation running through 2026 and beyond. Transparency obligations under Article 50 kicked in on August 2, 2026—effectively this week—with limited grace for systems placed before that date. Restrictions on real-time remote biometric identification in public spaces for law enforcement are already binding. Market surveillance authorities across EU member states were designated through 2025 and 2026.
The regulatory landscape, in short, favors approaches that preserve privacy by design, provide audit trails, and allow for human oversight. Companies marketing "AI that lives in your camera" will need to demonstrate how those systems comply with evolving rules in each jurisdiction. The companies that get this wrong won't survive contact with regulators.
Deployment Patterns, Emerging Fast

The use cases are becoming specific enough to pattern-match across verticals. Car washes are integrating license plate recognition with loyalty programs and point-of-sale discounts—DRB's SiteWatch documentation from 2024 through 2026 shows this is production reality, not concept. Culver's announced on May 1, 2026, that it would roll out computer vision from Berry AI to roughly 1,000 restaurants to monitor drive-thru throughput and order accuracy in real time.
Fitness centers and 24-hour gyms use existing cameras to detect tailgating at access-controlled entry points, integrating with systems like Kisi and Meraki, per case studies from 2025 and 2026. Industrial safety is another vertical gaining traction: Voxel, which raised a $44 million Series B in June 2025, publishes customer stories claiming outcomes like a 91% reduction in recordable incidents at one facility over two years.
Retail, logistics, and physical security span multiple sub-segments, each with slightly different needs. Novaira Insights noted in 2026 that the majority of new cameras now ship with embedded deep-learning analytics, though cloud-connected cameras still represent a single-digit percentage of the installed base. The gap between hardware capability and actual connected deployment? That's opportunity space for software platforms.
The Architecture Question
The prize is large: hundreds of millions of installed cameras, most still functioning as passive recorders, sitting in businesses that are just beginning to ask what automation those cameras could trigger. The technical path is clearing—edge chips powerful enough, bandwidth techniques validated in peer review, natural-language interfaces abstracting away the complexity of computer vision pipelines.
The question, though, is less whether cameras become agents and more which architecture wins. Will enterprises buy all-in-one edge appliances like OpenVector's Station, or adopt cloud platforms like Spot AI that backhaul video for centralized processing? Will they layer software onto cameras from Axis and Verkada that already embed analytics, or choose hybrid approaches that split inference between device and edge box?
OpenVector's bet is that foveated sampling efficiency, peer-reviewed at CVPR, creates a moat in bandwidth-constrained environments—IoT deployments, body-worn cameras, distributed facilities where connectivity is expensive or unreliable. Whether that technical edge translates to customer traction depends on factors the company hasn't yet disclosed publicly: pricing competitiveness at scale, integrations beyond the handful shown on the website, and the operational reliability required to displace manual workflows in car washes, restaurants, and warehouses.
The broader trend is unmistakable. Cameras already outnumber people in many commercial environments. The compute to make them intelligent is shipping. The software layer to make them autonomous is here, or close enough. What remains is the messy, expensive work of proving it works reliably enough that a business will trust a camera to make decisions a human used to make.
And the even harder work—building a company fast enough to capture market share before the window closes, before the established players adapt, before the next wave of well-funded startups arrives with their own CVPR papers and refundable deposits. Gizdov and Urlam have the research pedigree and, apparently, the conviction. Whether they have the timing is something only the market will reveal.
