Egocentric perspective of human hands manipulating mechanical components

Global Robotics Intelligence Platform

Teaching robots
the way humans
actually work.

Grip is the robot-ready data supply chain for physical AI. Capturing, enriching, and delivering high-quality egocentric human-task data at the scale foundation models need.

Egocentric captureRobot-ready schemasLeRobotMCAPROSDM0 quality standardsGlobal collection200k+ annotated hoursEgocentric captureRobot-ready schemasLeRobotMCAPROSDM0 quality standardsGlobal collection200k+ annotated hoursEgocentric captureRobot-ready schemasLeRobotMCAPROSDM0 quality standardsGlobal collection200k+ annotated hours

The gap

Robotics is missing 95% of the data it needs to scale.

Today's global data collection sits at roughly 1 million hours (about 10 human lifetimes). Frontier model teams estimate 100+ million hours are still needed to reach a true ChatGPT moment for physical AI.

At current rates, closing that data gap will take decades. Most efforts trade quality for volume, lean on teleoperation and synthetic data, or lack the production pipelines that model teams need at scale.

~1M
Hours collected globally
~100M
Hours still needed

The platform

An end-to-end data supply chain, from human hands to robot-readiness.

01

Capture

Real human egocentric recording in factories, homes, and workplaces — grounded in the environments robots will actually operate in.

02

Quality

Production-grade standards co-designed with frontier model teams through our DM0 partnership.

03

Enrich

Full-stack QC, annotation, and enrichment pipelines built for physical AI, not repurposed from web data.

04

Deliver

Robotics-native formats — LeRobot, MCAP, ROS, and custom schemas — ready to plug into model training runs.

The product

What the delivered dataset contains

The package is designed as an inspectable robotics dataset rather than a video-only demo. Every layer is synchronized around the same episode id.

01

Raw video

1920 x 1080 egocentric MP4

Primary observation stream for manipulation context.

02

Sensor JSON

IMU / capture metadata

Synchronizes motion signals with visual frames.

03

Depth video

512 x 288 dense preview

Human-readable inspection layer for scene geometry.

04

SLAM JSONL

Per-frame camera pose

Supports trajectory review, reconstruction, and point-cloud projection.

05

Depth samples

Frame-level .npy depth maps

Feeds metric inspection, 3D sampling, and volume estimation.

06

Hand labels

3D joints and transforms

Captures interaction geometry that plain RGB misses.

07

Annotation

Action, subtask, skill boundaries

Turns long videos into reusable robot-learning episodes.

The demo

Takeaway box packaging and goods placement

RGB

Video

Original raw egocentric video clip.

Depth

Video

Per-frame metric-style depth preview.

Hand Pose

Video

Rendered hand-pose labels over RGB.

The distribution

Multiple domains, scenes and task coverage areas.

The data sample format is consistent across domains, while the corpus emphasizes everyday unstructured spaces and repeatable fine-grained manipulation tasks.

Dataset
Scene and task coverage
Hours
Share
Home Operations
Covers real home spaces such as kitchens, living rooms, bedrooms, bathrooms, storage areas, entryways, laundry and cleaning areas, balconies, and similar domestic environments.
47,250
35%
Office Light Work
Covers desks, file management, desktop organization, printing, copying, binding, tool benches, and other lightweight office operations.
27,000
20%
Supermarket / Retail
Covers shelves, refrigerated cabinets, picking and replenishment, front-of-store display, backroom packing, large-scale stocking, organization, classification, and display tasks.
33,750
25%
Industrial / Semi-Structured
Covers tool stations, small-parts benches, label sorting, plug/socket alignment, tray or carton loading, and other high-precision fine manipulation tasks.
27,000
20%

The scale of our impact

200k+

Annotated hours of production-grade egocentric data already delivered

100k

Hours added every month across global collection locations

DM0

Active co-design partnership shaping quality and schema standards

The advantage

Built for the teams defining physical AI.

Grip is the only player combining large scale data inventory, model-team-defined data quality standards, and a global production pipeline, at competitive unit economics for data procurement budgets.

Aspect
Grip
Typical competitors
Data type
Egocentric, synchronized perception layers
Heavy robot teleop / synthetic
Quality standards
Co-designed with frontier model teams
Internal or research-oriented
Pipeline ownership
Full stack — hardware to delivery
Partial: collection or open datasets
Unit economics
Lower-cost, Emerging markets
Higher-cost, Western-centric
Production readiness
Native robotics schemas, ready inventory
Research sets or early pipelines

The next era of physical AI
will be built on Grip.

If you're training embodied models, we should talk.