Issue #047 — Weekly Trend

Share

Issue #047 — Weekly Trend

The TOPS Trap: What China's Jetson Challengers Still Can't Copy

By Ken ZHANG (Maze Intelligence)

The most important number in robotics edge compute this year is not a TOPS figure. It is a price: $1,799, roughly double the $899 list price of a Jetson AGX-class module before the current supply squeeze (EET-China, 2026-08). For every robot integrator in Shenzhen, Shanghai, and Singapore, the question has quietly flipped from "should we qualify a domestic module?" to "which one, and what will the switch actually cost?"

The scoreboard says the gap is closed. The honest answer is that the scoreboard is measuring the wrong thing.

The paper war is over

Walk down the 2026 domestic lineup and the TOPS race looks, on paper, finished:

  • Racobit's Shanhai T100 (built on an undisclosed domestic GPU) claims 240 TOPS INT8 with 32GB of memory — squarely in AGX Orin territory at 275 TOPS (NVIDIA).
  • Moore Threads' E300 module, built on its in-house "Yangtze" SoC, delivers 50 TOPS of dense INT8 at under 45W — and up to 64GB of unified LPDDR5X (Moore Threads documentation).
  • Horizon's S100P puts 128 TOPS on-device and is already shipping inside quadrupeds (Xinhua, 2025-12).
  • Sophgo's SM9 is pin-compatible with Jetson carrier designs — a socket-swap for cost-down programs.
  • Huawei's Atlas 200I A2 (8/20 TOPS INT8) and Rockchip's RK3588 (6 TOPS) anchor the low end, where the unit volume actually lives.

Two years ago, this column of numbers did not exist. That is real industrial progress, and anyone dismissing it as subsidy cosplay has not watched a Chinese toolchain mature at the current pace.

But TOPS parity is not Jetson parity — for three reasons the datasheets underplay.

Reason one: the memory wall

A modern robot policy stack — a quantized VLA or VLM, perception models, and the orchestration layer on top — is memory-bound long before it is compute-bound. Capacity decides what fits on the robot. Bandwidth decides how fast it thinks.

NVIDIA's quiet moat is that AGX Orin 64GB pairs that capacity with 204.8 GB/s of bandwidth (NVIDIA product specifications). Among domestic modules we could find, the best published figure is Moore Threads' 136 GB/s — about two-thirds of NVIDIA's — and the Shanhai T100 does not publicly disclose bandwidth at all. Most of the rest sit at 16GB of memory or less, which caps them at perception workloads no matter what their TOPS rating says.

In other words: domestic silicon can now count the operations. It cannot yet feed the model.

Reason two: CUDA is shallow at inference, deep everywhere else

The compatibility story has genuinely improved. Moore Threads' MUSA SDK now tracks CUDA 12.8 with 761 compatible interfaces, full coverage of the core math libraries, and vLLM as an official backend — porting an inference stack is now a weeks-scale project, not a rewrite (Moore Threads, 2026).

But inference was always the easy layer. The training-and-simulation stack — Isaac Lab, and the sim-to-real toolchain most embodied-AI teams now build their pipelines on — remains CUDA-locked, a point we pressed in Issue #041. You can port your robot's runtime. You cannot port the ecosystem your engineers were trained on.

Reason three: the undisclosed node

The most aggressive challenger, the Shanhai T100, will not say whose GPU is inside it. From a supply-chain standpoint, that is an unverifiable node in your BOM: you cannot audit what you cannot name, and "100% domestic" is a claim, not a provenance. Any serious qualification program should treat supplier disclosure as a gate, not a footnote.

The buyer's matrix, condensed

For teams making this call in Q4 2026:

  • State-affiliated tenders: Ascend Atlas 200I A2. Weakest TOPS on this list, strongest ecosystem — CANN, and Advantech-class integrators already shipping Atlas-200I-A2 edge platforms — plus procurement scoring that rewards it.
  • Cost-down on existing Jetson designs: Sophgo SM9. Pin-compatibility first; revalidate everything anyway.
  • High-TOPS new designs: Shanhai T100 — but only after two gates: supplier disclosure written into the contract, and a benchmark on your models, not the vendor's.
  • On-robot VLA/VLM development: still AGX Orin 64GB. Nothing domestic we could find both fits the memory envelope and has volume field mileage. Watch the E300's first design wins.

Three falsifiable bets

One: by mid-2027, domestic compute modules appear as a scored line item — not a bonus — in the majority of Chinese state-affiliated robotics tenders. Two: no domestic module ships in volume before then with published memory bandwidth above 136 GB/s. Three: the first tier-1 robot OEM design win on E300-class memory (64GB unified, domestic silicon) lands within 12 months. If any of the three breaks, the correction runs here.

The irony is that NVIDIA's price hike has done more for Chinese edge silicon than any industrial policy: it converted the pitch from nationalism to economics. TOPS parity got domestic vendors to the table. The memory wall and the ecosystem lock decide who actually eats.


Sources: NVIDIA Jetson AGX Orin product specifications (nvidia.com); Moore Threads E300 product documentation and 2026 launch materials (docs.mthreads.com); EET-China (2026-08) on Shanhai T100 specs and Jetson supply/pricing (eet-china.com/mp/a510599.html); Huawei Ascend community Atlas 200I A2 specifications (hiascend.com); Xinhua (2025-12-15) on Horizon S100P deployments. Pricing figures are trade-press reports; verify with distributors before BOM commitment.