Skip to content
Artwork for Embodied AI 101
TechnologyScience

Embodied AI 101

Shaoqing Tan

Stay in the loop on research in AI and physical intelligence.

Play
  • 55 episodes
  • Avg 28 min
  • English
Counted on this page — what you have heard stays on this device, so it is not something the list can be paged by.
  • Yesterday · 28 min

    Tune Slowly, Control Quickly: Learning a Better Robot Navigation Stack

    Autonomous navigation in complex, unstructured environments poses a significant challenge, with traditional planners lacking adaptability and end-to-end learning methods hindered by data dependency or training instability. This paper proposes a hierarchic

    • Transcript
  • Thursday · 29 min

    THAW-VLA: World-Model Features Without World-Model Latency

    This work distills frozen world-model representations into compact VLAs via a single feature-alignment loss at training time, with the teacher cached once and absent at inference; a 0.8B student achieves 97.9% on LIBERO and lifts RoboCasa-GR1 humanoid per

    • Transcript
  • Thursday · 28 min

    Joga: Why a Soccer Humanoid Needs to Control Its Gaze

    Joga demonstrates agile humanoid soccer on a Unitree G1 using an actuated neck and residual models to close the sim-to-real gap for both perception and actuation, enabling dribbling, tight turns, and receive→dribble→shoot sequences with fully onboard perc

    • Transcript
  • Wednesday · 29 min

    AthenaZero: Why Dynamic Manipulation Starts With Lower Inertia

    AthenaZero is a bimanual manipulator designed to minimize inertia without compromising control authority. By using quasidirect drive actuation and transmission remotization techniques, the system achieves an effective end-point mass comparable to state-of

    • Transcript
  • Wednesday · 29 min

    FROA-Drive and the Hard Part of Targeted Policy Repair

    Vision–Language–Action (VLA) models have shown strong potential for end-to-end autonomous driving, yet their post-training commonly relies on expensive simulator interaction or global policy updates. FROA-Drive introduces a failure-routed offline adaptati

    • Transcript
  • Tuesday · 28 min

    Teach the Robot Where to Act—Then Make It Fast

    We present a framework for assistive robot manipulation that addresses two fundamental challenges: efficient adaptation of large-scale models for scene affordance understanding and effective learning of robot actions by grounding...

    • Transcript
  • Tuesday · 28 min

    HEAR: The Sound Your Robot Missed Between Actions

    Humans and animals use sound as a crucial cue for interacting with the physical world, as acoustic events can reveal contact, completion, hidden contents, or process state. Embodied agents should leverage sound for manipulation tasks. This paper presents

    • Transcript
  • Sunday · 29 min

    VS-Splat: Learning Where 3D Detail Belongs

    A new feed-forward Gaussian Splatting method that selectively processes voxels for efficient 3D scene reconstruction, targeting real-time perception stacks for embodied agents. Advances the efficiency frontier of 3D scene representation without per-scene

    • Transcript
  • Sunday · 13 min

    Spatial Search Comes to Quest — Beyond Pretty Gaussian Splats

    Adapts a training-free Spatial Search prototype to Meta Quest's spatial data pipeline, enabling natural-language search over Gaussian Splatting and mesh scenes in real time on-device. Demonstrates practical deployment of 3D spatial intelligence for consum

    • Transcript
  • September 19 · 30 min

    Occamy-1.0: Teaching a 35B Agent to Finish the Job

    Introduces a 35B cost-efficient model positioned at the Pareto frontier for collaborative long-horizon agent tasks, achieving performance competitive with significantly larger systems. Advances the capability of embodied and agentic AI to handle extended

    • Transcript
  • September 18 · 12 min

    Helix 2.5: The House Changed. The Policy Didn't.

    Figure's Helix 2.5 humanoid robot enters 30 unseen Bay Area rental homes with no additional training and performs useful whole-body tasks for 4+ hours, demonstrating strong zero-shot generalization across diverse real-world environments. This is a notable

    • Transcript
  • September 16 · 12 min

    The Hard Part Was the Stack: PI's Robot Goes to Work

    Physical Intelligence's mobile π robot runs fully autonomous for hours in a real production environment stacking boxes at Dandelion Chocolate, exposing generalization and reliability gaps between lab demos and production utility. Represents a significant

    • Transcript
  • September 16 · 13 min

    ZDTaichu5.0-9B: Better Spatial Reasoning, an Unfinished Edge Story

    ZDTaichu5.0-9B is a 9B edge-deployable multimodal model built on a Qwen3.5 backbone with C-RADIOv4 vision encoder that leads open sub-10B VLMs on spatial-reasoning benchmarks including ViewSpatial and MMSI-Bench, while supporting embodied AI and tool-use

    • Transcript
  • September 15 · 28 min

    Light-Loco-Parkour: Learning the Skill Is Only Half the Problem

    Light-Loco-Parkour presents a multi-skill distillation approach for learning perceptive whole-body parkour and locomotion skills deployable on real robots, enabling versatile agile movement across diverse terrains. The method advances embodied locomotion

    • Transcript
  • September 15 · 29 min

    OM-1: Human Skills, Many Bodies

    OM-1 is a robot foundation model trained exclusively on human manipulation data (no teleoperation or robot demonstrations) that achieves near human-level dexterity and zero-shot generalization across tabletop arms, industrial arms, and humanoids, includin

    • Transcript
  • September 14 · 12 min

    Show-Harness Gives VLMs a Robot Keyboard

    Uses discrete semantic actions paired with embodiment-specific interpreters to turn any pretrained VLM into a robot controller without a dedicated policy network or additional pretraining. Demonstrates strong zero-shot generalization across diverse tasks

    • Transcript
  • September 14 · 29 min

    BridgeVLA++: Keep the Geometry, Give the Robot a Memory

    Renders point clouds as three orthographic 2D views fed into a pretrained PaliGemma VLM, predicting heatmaps to recover 6-DoF actions with temporal/spatial memory modules. Achieves 95.4% success on 13 real Franka tasks with only 3 demos per task, outperfo

    • Transcript
  • September 13 · 28 min

    The Sim-to-Real Gap Inside the Gearbox

    Reinforcement learning (RL) has become a powerful tool for quadrupedal locomotion, and a sim-to-real approach is widely adopted to avoid hardware damage during training. However, the sim-to-real gap remains a significant challenge.

    • Transcript
  • September 13 · 28 min

    A Stable Grasp Isn't Enough: Giving 6-DoF Grasping a Purpose

    Robotic grasping in cluttered scenes is dominated by methods that optimise either stability or a predefined task category, leaving a gap between 'can this grasp hold the object' and 'can this grasp serve the task'. This paper bridges that gap using founda

    • Transcript
  • September 13 · 29 min

    HiDream-O1-Embodied: A World Model Must Earn Its Actions

    HiDream-O1-Embodied is a unified native architecture that ingests text, images, and video and directly outputs physical actions, positioning itself as a step beyond passive world modeling toward active embodied interaction. The architecture aims to close

    • Transcript
Showing 1–20 of 55 episodes