Find, compare and buy robots from top manufacturers

Robbyant Launches LingBot-VA 2.0 for Real-Time Robot Control

Robbyant’s new world-action model slashes execution latency to 142ms, redefining how industrial robots and AMRs operate in the global marketplace.

Image Credits:
Robbyant

Harper Whitmore

Robotics News Reporter

Watch a state-of-the-art robot try to clear a messy dinner table, and you will inevitably witness a frustrating micro-hesitation. The machine reaches for a glass, stops, computes the next frame of its reality, and only then executes the shift. In the lexicon of robotics, this is the latency penalty: a cognitive lag born from forcing physical machines to use brains trained on digital data. For years, robots have been forced to react to the world rather than move through it fluidly.

That paradigm changed this week. Robbyant, an embodied AI company operating within Ant Group, announced the release of LingBot-VA 2.0. Rather than fine-tuning an existing video generator to control hardware, Robbyant built LingBot-VA 2.0 from scratch as an embodied-native world-action model. The system introduces a mechanism called Foresight Reasoning, allowing a robot to internally simulate and plan its next physical maneuver while its mechanical limbs are still completing the current one.

The Myth of the Retrofitted Video Generator

For the past two years, the robotics industry has leaned heavily on a clever shortcut: borrowing generative video models trained on vast troves of internet media and bolting a robotic action policy onto the backend. The logic seemed sound. If a model understands how pixels shift when a ball bounces or a cup falls, it must understand the physics of the real world.

Yet, in practice, this approach yielded brittle automation. Content-generation models care about visual plausibility, not physical truth or real-time execution boundaries.

“Robbyant will continue to explore new limits in embodied intelligence while accelerating the development of an open technology and application ecosystem to expedite robot deployment in industrial and real-world scenarios,” said Zhu Xing, CEO of Robbyant.

Instead of predicting pixels for digital consumption, LingBot-VA 2.0 unifies vision and action tokens into a single semantic latent space. In plain terms, the robot does not translate its visual surroundings into text commands and then into motor outputs. It maps what it sees and what it does onto the exact same coordinate system. This architecture allows the model to absorb unlabeled internet video data while keeping the ultimate output strictly focused on physical control.

Breaking the 900-Millisecond Barrier

The true breakthrough of LingBot-VA 2.0 is not merely conceptual; it is an engineering feat of raw speed. Traditional autoregressive models are notoriously heavy, requiring massive computational time to generate the next sequence of actions.

Previous iterations suffered from a crippling 927 ms execution latency per chunk, an eternity when a robot is attempting to catch a moving object or assemble delicate electronics. Robbyant tackled this bottleneck by introducing a four-stage inference stack that combines FP8 TensorRT compilation and long-horizon attention optimization.

The result is a dramatic compression of computational lag:

Performance MetricTraditional Baselines / Version 1.0LingBot-VA 2.0
Per-Chunk Latency927 ms142 ms
Control Frequency35 Hz225 Hz
Single-GPU Inference—150 Hz
Bimanual Task SuccessStandard Baseline93.6%

By slashing per-chunk latency to 142 milliseconds, the system increases asynchronous control from a stuttering 35 Hz to a fluid 225 Hz. This capability relies heavily on a sparse Mixture of Experts (MoE) architecture. The model hosts 128 specialized experts, routing the top 8 for any given token. Out of the model’s total 15.3 billion parameters, only roughly 2.5 billion fire at any single moment. This design expands the robot’s capacity to understand complex environments without increasing the computational tax on real-time execution.

The Few-Shot Leap on the Factory Floor

Beyond pure execution speed, the new architecture addresses the data scarcity problem that has historically restricted industrial robotics. Typically, teaching a robot a new task, like sorting conveyor-belt items or tidying an unfamiliar desk, demanded hundreds of hours of precise, scripted data or simulated environments.

LingBot-VA 2.0 bypasses this requirement via robust in-context learning. Because the model relies on a strict causal architecture that understands how an action alters an environment, it requires as few as 20 human demonstrations to master an entirely new physical task. Crucially, it does so without requiring a single parameter update or traditional fine-tuning. An operator can record a single video demonstration of a task, and the model can decode the necessary visual-action pathways directly.

In real-world testing environments, Robbyant demonstrated this across diverse profiles:

  • High-speed air hockey matches requiring instantaneous micro-adjustments.
  • Fragile semiconductor chip picking where over-correction means destruction.
  • Long-horizon bimanual manipulations evaluated on the RoboTwin 2.0 simulation benchmark.

Sourcing the Brains for Tomorrow’s Hardware

For industrial buyers and automation engineers navigating global marketplaces like Anton Robots to source state-of-the-art collaborative robotic arms or autonomous mobile robots (AMRs), this shift completely alters the calculus of deployment. Traditionally, buying the physical hardware was only half the battle; the hidden, often prohibitive costs lay in integration, custom programming, and environment mapping.

Platforms like Anton Robots, which aggregate a highly diverse industrial robot ecosystem, are poised to become the frontline where these advanced AI brains meet physical chassis. When deployment times drop from months to hours via just 20 visual demonstrations, the economic barrier to entry for factory automation collapses entirely.

The rollout of LingBot-VA 2.0 represents the capstone of a broader, coordinated push by Robbyant, which quietly unveiled six distinct models during its recent launch week. Together, these systems, ranging from the LingBot-Map 3D spatial streaming framework to LingBot-Depth, create a full-stack architecture for physical AI.

The broader robotics landscape is moving rapidly, with players chasing the goal of general-purpose physical agents. However, Robbyant’s focus on slashing the processing delay between a robot’s perception and its execution draws a clear line in the sand. For decades, the promise of automation was held back by machines that could either think deeply or act quickly, but rarely both at once. By forcing its models to imagine the future while executing the present, LingBot-VA 2.0 suggests that the future of robotics won’t belong to the machines that react fastest, but to the ones that know how to anticipate.

Related Insights

News

Chinese Humanoid Robot Breaks Usain Bolt’s 100-Meter Sprint World Record

News

Honor’s Flash Humanoid Completes Autonomous Half Marathon in Under 51 Minutes

News

Technical Hitches Disrupt Live Demonstrations at World Robot Conference

News

World Robot Conference 2026 Kicks Off in Beijing with 3,000 Autonomous Products

Looking for a Robot for Your Business?

Find the right robot based on your application, industry and requirements, or explore and compare available models.

Browse Robots

GET STARTED

Find My Robot

Compare Robots

BROWSE BY TYPE

Humanoid Robots

Robot Dogs

Robotic Arms

Cobots

AMR Robots

AGV Robots

Service Robots

Companion Robots

BROWSE BY APPLICATION

Material Handling

Palletizing

Pick and Place

Welding

Inspection

Cleaning

Delivery

Security

BROWSE BY INDUSTRY

Manufacturing

Warehousing

Medical

Restaurants

Agriculture

Construction

Retail

Education