We just released an end-to-end system that sets a new state of the art in egocentric video understanding, generating fine-grained action labels from raw robot and human videos while outperforming Gemini, GPT, Claude, and other leading models.
We're open-sourcing our infra with 10M+ frames of dataset!
We're releasing Stera, an open-source infra that turns an off-the-shelf device in your pocket into a high-fidelity multimodal data pipeline. It's built around four layers. Capture → Process → Evaluate → Export.
Stera Capture removes the need for bespoke/gated hardware and runs on an off-the-shelf iPhone. It fuses together synchronized RGB, IMU, Lidar-guided depth, and 6-DoF pose out of the box from ARKit and exports them to a raw MCAP file.