Multimodal Data

Multimodal training data from real work: annotated screen trajectories, voice interactions, documents, and video demonstrations of professional workflows, captured with full consent.

Collage of photographs, waveform strips, and typed sheets pinned to paper

One expert episode, every modality

When an expert walks a task in our recorded sandbox, the capture is total: what the screen showed, what the hands did, what was said, what documents were produced, and whether the outcome verified. That single aligned episode is worth more than the same hours of scraped video, because every frame is grounded in a task with an explicit notion of done.

The modality mix follows the work

Some professions live in the browser, some on the phone, some in a stack of PDFs. Dispatch runs on load boards and live calls. Billing runs on portals and worklists. Bookkeeping runs on ledgers and statements. We source the capture format from how the work actually happens rather than forcing everything into screenshots.

Built to extend

The capture pipeline is pluggable by design. Teams working on world models and embodied research have asked the same question in different words: can your experts film what they do all day. Yes, and the footage arrives attached to task structure, rubrics, and verified outcomes rather than as raw video.

What ships in the box

Screen trajectories

Full desktop and browser recordings of expert work, aligned with the action log, clicks, keystrokes, and state changes, frame by frame.

Voice at work

Spoken workflows and work calls, transcribed and aligned to outcomes, from professions where the voice channel is the job.

Documents as they exist

The spreadsheets, ledgers, forms, and files real workflows produce and consume, with the messy formatting intact because the mess is signal.

Video demonstrations

Experts recording their physical and on-screen day-to-day, a pluggable capture format that extends to embodied and world-model research.

Aligned across modalities

One episode yields synchronized screen, action, audio, document, and outcome data, not separate corpora stitched after the fact.

Consent and compensation

Every contributor is told plainly their data trains and evaluates AI systems, and is paid for every stage of participation.

See it before you buy it.

Sample task packets and environment access for evaluation. Tell us the capability you care about and we will send the relevant cut.