Jensen Huang stood at NVIDIA's GTC Washington DC keynote in June and delivered a succinct frame: "Let's talk about physical AI." By that point, the moment had essentially arrived. Vision-language models had broken the sub-second barrier in production, and security cameras are increasingly being integrated with AI for real-time image analysis—capabilities that remain an ongoing development across more than a billion cameras installed worldwide, with systems moving toward answering questions about what they see, in real time, without sending the footage to the cloud.
The shift registered most visibly in MLPerf Inference v6.1, released September 17, which showed the biggest performance gains in VLM workloads since the benchmark category launched six months earlier. Miro Hodak and Frank Han, who chair MLCommons, described the results as an inflection toward "agentic and end-to-end benchmarking." What had once been the domain of computer vision—detect this object, flag that anomaly—now involved multimodal reasoning over video streams, natural-language queries, and autonomous responses to safety violations or manufacturing defects.
The stakes are considerable. AP reported in March that more than one billion security cameras sit installed globally. Omdia estimated the video surveillance market at $27 billion in 2025. Companies deploying vision AI expect far more than object detection. They want systems that can field natural-language queries over live footage, reason through complex manufacturing defects, and act on safety violations without human intervention.
Thirty Submitters, 120 Systems, and a Record Field
According to MLPerf chairs, v6.1 drew thirty submitters and 120 systems this September, the broadest field to date and what the chairs called "the broadest field of submitters we have ever seen." The round introduced an Interactive scenario for VLM workloads based on Qwen3, measuring how quickly systems could process image-plus-text prompts and return answers. One submission ran on 512 accelerators, set a record, and delivered roughly 5.8 million tokens per second on GPT-OSS-120B in offline mode—a signal, perhaps, of the larger VLM and agent workloads coming.
Enterprise deployments now span workplace safety, manufacturing quality control, smart city traffic management, and port logistics. Voxel, a workplace safety vision AI startup that announced $44 million in Series B funding on June 3, 2025, reported that customers including Autokiniton deployed its system in roughly one week and saw up to 80 percent reductions in high-risk behaviors and 91 percent reductions in recordable incidents. NSG Group rolled out Voxel globally in October 2025.
Pegatron scaled visual AI and digital twins for factory inspection workflows using NVIDIA Metropolis, according to a case study published in 2026. The Port of Barcelona deployed more than 540 Axis cameras running on-device analytics to optimize logistics and crowd management while sharing only metadata for privacy, Axis said in September. Cities including Raleigh and San Jose piloted video analytics agents built on NVIDIA's DeepStream for traffic analysis, NVIDIA case studies show.
Three Breakthroughs, One Latency Collapse

The technical story behind the sub-second barrier involves three breakthroughs: disaggregated inference pipelines, aggressive KV-cache compression, and hardware-aware attention kernels.
Encode-Prefill-Decode disaggregation splits vision encoding from language prefill and decode, NVIDIA explained in a September technical blog introducing its Dynamo framework. The strategy works when vision processing dominates total compute or when output responses run long. SGLang adopted EPD in its multimodal roadmap during the third quarter of 2026. "The benefits depend on vision load and response length," NVIDIA wrote, which is another way of saying the approach doesn't help every workload.
Researchers cut memory and bandwidth bottlenecks by exploiting VLM-specific sparsity patterns. ZipVL, presented at ICCV 2025, reported 2.3× prefill speedups and 2.8× decode speedups on LLaVA-Next-13B with roughly 0.5 percent accuracy loss through dynamic token sparsity. Q-Cache, published at AAAI 2026, found visual attention mattered in fewer than 50 percent of decode layers. VL-Cache, introduced at ICLR 2025, compressed KV cache by identifying unique visual versus text sparsity in prefill and decoding.
FlashAttention-3, widely deployed through 2026 in vLLM and TensorRT-LLM, combined FP8 quantization with hardware-optimized kernels to increase throughput. NVIDIA's TensorRT-LLM added vision encoders with tensor and context parallelism plus FP8 support for Qwen2-VL in its September release notes. vLLM documented multimodal coverage spanning LLaVA family models, Phi-3.5-Vision, InternVL 3.x, and Qwen-VL variants, with an FP8 KV cache option shipping in April 2026.
CV-CUDA offloaded image and video preprocessing to the GPU, delivering up to 49× end-to-end speedups in video segmentation pipelines, NVIDIA said. The tool addresses a recurring bottleneck: community benchmarks from August 2026 showed multi-camera RTSP decoding and zero-copy NVDEC workflows often dominate total latency over the model itself in real deployments.
SemiAnalysis InferenceX, which published measured cost-per-token and throughput data from June through September, showed steep cost and throughput improvements generation-to-generation across NVIDIA's GB200, GB300, and Vera Rubin NVL72 rack-scale systems when using disaggregated pipelines.
Edge Deployments Run on Watts, Not Kilowatts

Edge deployments bring VLMs onto cameras and robotics controllers with power budgets measured in watts, not kilowatts. Ambarella launched its CV7 SoC in January 2026, enabling transformer-based AI and VLMs to run on-camera for enterprise security. The chip maker's Model Garden listed curated VLMs including InternViT variants and LLaVA-OneVision-Qwen2-7B for CV7-class devices. Verkada announced Ambarella CV7x integration for on-camera analytics including person-of-interest alerts and license-plate recognition in February 2026.
NVIDIA's TensorRT Edge-LLM, introduced in January for automotive and robotics, targets deterministic low-latency embedded inference with EAGLE-3 speculative decoding and NVFP4 quantization. Jetson benchmarks published in 2026 showed Qwen2.5-VL 3B running on Jetson AGX Orin. A TensorRT Edge-LLM user guide listed Qwen3-VL-2B performance in INT4 and FP16 on Orin.
EdgeFM, an agent-driven cross-platform inference framework published in April 2026, reported up to 1.49× speedups over TensorRT-Edge-LLM on NVIDIA Orin in the authors' tests. TurboVLA, published in July, achieved 32 Hz on an RTX 4090 with less than 1 GB VRAM for robotics vision-language-action pipelines.
OpenVINO 2026.4, released mid-September, added VLM-specific controls including separate configurations for language models and vision encoders, visual-token support on Intel Core Ultra Series 3 NPUs, and preview support for DFlash acceleration on Qwen.
Cisco Meraki MV and Axis edge cameras run on-device analytics today for people counting and object detection, with metadata-only sharing patterns documented in 2026 datasheets. Vendors position the architectures as privacy-preserving while experimentally enabling bring-your-own VLM models.
T-Mobile collaborated with NVIDIA on VSS Blueprint v3 to integrate physical AI applications on AI-RAN-ready infrastructure, the companies announced March 16. The City of San Jose began early assessment of the platform for smart city video analytics.
Agentic AI Meets Physical Constraints

MLPerf chairs said they expect broader agentic and end-to-end benchmarking in future rounds as VLM participation climbs. Sequoia Capital wrote in a September 2 thesis that 2026 and 2027 applications would become "doers"—agentic AI moving from talk to action, including physical-world tasks aided by vision. McKinsey linked operational excellence with faster AI scaling in manufacturing and logistics in a May article published in July.
IDC reported a pilot-to-production shift in edge AI during the first half of 2026, according to an August perspective document. Stanford's AI Index 2026 showed increased "agent" use in operations and IT, underscoring corporate adoption momentum.
Regulatory and privacy pressures complicate the outlook, though. The EU AI Act, finalized in June 2024 with enforcement phasing through 2026 and beyond, prohibits real-time remote biometric identification in public spaces for law enforcement except under strict exceptions. Transparency obligations under Article 50 took effect August 2, 2026. The FTC issued a proposed policy statement on deception regarding AI system accuracy in 2026 and maintains AI use inventories across agencies.
Public backlash intensified in the United States. Senate Judiciary opened an inquiry into Flock Safety's license-plate-recognition and AI camera practices on August 26, Axios reported. Flock CEO Garrett Langley apologized "for instances where our data was misused by law enforcement officers" in an August 14 CBS News interview and announced platform changes. Regional pushback emerged across the D.C. metro area in September.
The 2021 Verkada breach remains a reference point for cloud-managed camera security. Founders deploying vision AI should monitor MLPerf submissions for inference cost trends, evaluate EPD disaggregation against their median video-to-response ratios, and architect for metadata-only sharing where regulations or customer expectations demand it.
Companies building on more than a billion installed cameras now have the inference stack to make them reason in real time. Whether they should remains the harder question.
