
Technology NewsJuly 20, 2026
Cosmos 3 and Qwen 3.5, measured at the edge.
YUAN benchmarks three reasoning vision-language models on the NVIDIA Jetson T5000 — and finds a clear trade between semantic depth and raw throughput across 49 concurrent camera streams.

AI is leaving the data centre for the physical world, and the next generation of edge video AI is being asked to do more than detect objects: it must reason, in multiple modalities, at low latency, across dozens of streams at once. To map that territory, YUAN ran in-depth performance benchmarks of three next-generation reasoning vision-language models — NVIDIA Cosmos3-Nano, Cosmos3-Edge, and Qwen3.5-9B — on the NVIDIA Jetson T5000 platform. NVIDIA Cosmos 3 is a world foundation model that can serve either as a perception-to-action model or as a reasoning VLM; it is the reasoning role YUAN put under test.
Accuracy against throughput
The headline result is a throughput story. At 49 concurrent streams, Cosmos3-Nano sustained 525.01 total tokens per second — more than twice the 215.37 tokens per second of Qwen3.5-9B — while giving up only about one point of accuracy: 96.31 per cent against 97.44 per cent. Both figures come from vLLM with NVFP4 quantization, YUAN's production-optimised configuration. For high-density, multi-camera edge deployments, that is a strong trade.
Qwen3.5-9B remains the model to pick when semantic understanding matters most. With 97.44 per cent accuracy and an F1 score of 91.15, it offers the best overall balance of precision and recall in the field — the premier choice for applications that depend on reliable contextual reasoning. Cosmos3-Nano, at 96.31 per cent accuracy and an F1 of 88.14, is the throughput engine, built for wringing the most streams out of a single module.
An equal footing for the whole family
Cosmos3-Edge — the lightest member of the Cosmos 3 family, designed for the most constrained devices — does not yet support NVFP4 quantization or the vLLM runtime. So YUAN re-ran all three models in full FP32 precision on the Hugging Face Transformers runtime, producing a clean baseline independent of quantization or serving-framework optimisations. On that footing, F1 score — the fairer cross-model metric when class imbalance can skew raw accuracy — scales with model size: Qwen3.5-9B at 90.32, Cosmos3-Nano at 88.32, Cosmos3-Edge at 87.40.
At single-stream throughput the order flips, as expected for an unquantized, unbatched baseline: Cosmos3-Edge leads at 7.37 tokens per second, ahead of Cosmos3-Nano at 3.83 and Qwen3.5-9B at 2.95. Where every millisecond and milliwatt counts — a single camera, a power-constrained cabinet — Cosmos3-Edge is the natural fit. For multi-camera production workloads, Cosmos3-Nano under NVFP4 and vLLM remains YUAN's recommended path.
From video analytics to physical AI
Traditional video AI is bound to predefined object classes. Pairing YUAN's edge reasoning stack with multimodal VLMs lets an operator simply ask the system a question — has a traffic accident occurred on this road, is there an anomaly in the monitored area — and get an answer grounded in the live scene. By combining the Jetson T5000's compute with YUAN's video capture, multi-camera integration, SmartNVR, and SmartVMS technologies, developers can run scene understanding and event reasoning entirely on-site — keeping data private while moving video analytics towards physical AI and agentic applications that decide for themselves.
Related
More from the newsroom.



Talk to the team behind the story.
Technical questions, review hardware, and interview requests go through YUAN communications.
Contact PR