High-quality STEM problems
Original master's-level and above problems in maths, physics, chemistry and biology, each with a detailed worked solution.
Collov Data turns real expert work, real spaces and real robots into datasets frontier labs can train on today — expert-built post-training sets, parsed interiors and 2.3 million real-robot episodes.









Unretouched first frames from delivered episodes. Every frame maps to a specific episode, robot serial number, capture configuration and MCAP recording you can replay before you buy.
Off-the-shelf reasoning, code, agentic, expert-domain and pretraining sets for language and multimodal models.
See the catalogSpatial dataReal interiorsParsed scenes, multi-step edit trajectories and spatial evals from Collov's production visual-agent pipeline.
See spatial dataPhysical AI data2,341,350Physics-verified real-robot manipulation episodes across 8 commercial domains and 34 robot platforms.
See robot dataModel training data
Original, expert-written data for reasoning, code, agents and professional domains — ready to ship today and already delivered to tier-one model labs. Download a sample before you talk to us.
Original master's-level and above problems in maths, physics, chemistry and biology, each with a detailed worked solution.
Medicine, finance, law and computer science at senior-undergraduate to PhD level, with professional answers to real-world problems.
Expert-built challenges with Docker environments, attachments, verifiable grading and an oracle answer for every item.
Open a row for details and a sample.
Original problems in mathematics, physics, chemistry and biology at master's level and above, each with a detailed worked solution. Built for deep-reasoning training.
View sample dataHigh-quality questions from medicine, finance, law and computer science, from senior undergraduate to master's and PhD level. Covers real-world problems with professional answers to strengthen understanding and generation in complex professional domains.
View sample dataDetailed diagnoses of the logical and factual errors in model outputs. Trains a model to review its own work — spotting and fixing factual errors, contradictions and reasoning leaps — for more rigorous, reliable answers.
View sample dataOriginal competitive-programming problems with sample input and output, editorial solutions and time and space complexity analysis. Covers classic algorithms and data structures for ICPC, LeetCode and Codeforces-style problem solving.
View sample dataPretraining corpora for a range of needs, covering text, images and video.
Request a sampleOriginal illustrated problems in physics, chemistry and biology from senior undergraduate to master's level and above, with detailed worked solutions, for multimodal deep-reasoning training.
View sample dataAgentic task samples drawn from real, high-value knowledge work. Each includes a task description, the related files and grading rubrics.
Request a sampleMulti-turn instructions from real overseas users to build, revise and maintain high-quality websites over time, each reviewed and scored by design experts. Suited to high-quality trajectory generation and reinforcement learning for code generation and agent behaviour.
Request a sampleBuilt by security professionals to improve model cyber capability. Each item includes the challenge, required environment and attachments, the Docker environments needed to solve it, a verifiable evaluation method and an oracle answer. Every challenge is currently unsolved by frontier models.
Request a sampleHigh-quality questions in medicine, finance and law with point-by-point grading rubrics, from senior undergraduate to master's and PhD level. Covers real-world questions and professional scoring points to improve answers in these domains.
Request a sampleSpatial data
Every run of Collov's visual agent produces structured spatial data as a by-product. Collov Data packages it for frontier labs, robotics teams and enterprises whose AI has to work indoors — with people reviewing along the way.
The loop that runs Collov's products — perceive the scene, plan, act, check the result, learn — writes a structured record at every pass. That record is the dataset.
Real interiors broken down into objects, materials, geometry and the spatial relationships between them — structured, queryable scene state rather than flat labels.
Complete multi-step edit histories from production traffic: what was asked, what changed, what was corrected. Not synthetic one-shot before/after pairs.
Benchmarks that check whether a model respects real-world constraints — depth, planes, occlusion, scale and placement — built on the same approach as Collov's Spatial Arena.
Humans stay in the loop: reviewers check and correct agent output before it becomes training data.
Physical AI data
A production data supply, not a research dump. Licensed, deduplicated and physics-verified episodes captured on real hardware, shipped as MCAP and LeRobot v3.0 — and expandable to your task list on a six-week capture cycle.
Compute scaled. Architectures scaled. Physical interaction data did not. Passive footage is missing the four things a manipulation policy needs most.
Video shows what moved, never the force that moved it. Grasp stability can't be learned from pixels alone.
No joint targets, end-effector commands or gripper state — nothing for a policy to imitate.
Annotators can't check momentum or kinematic limits, so noise and teleop artefacts slip into training.
Single-robot corpora overfit one kinematic chain and fail on a new body.
The robot's own state at up to 1 kHz, the commanded action, the resulting contact and the sensed outcome — time-synchronised in one MCAP episode.
Every domain below is live, delivered volume — not a roadmap. Select one to see its scenarios.
Clothing foldingThe deepest domestic manipulation corpus in the catalog: kitchens, bathrooms, bedrooms, living rooms and entryways, captured in real furnished apartments rather than staged rigs.
Delivered episodes by scenario
Top 17 of 250 scenarios. 233 more in the library, including kitchen cleaning, cooktop cleaning, mixed-clutter sorting and food handling.
Wheeled mobile manipulators, full-size humanoids and force-controlled arms — the range of bodies a generalist policy needs, in one licence.
Delivered hours, top seven platforms
Fleet composition
Parallel grippers, five-finger dexterous hands and vacuum tooling, with 6-axis wrist force/torque sensing.
Signal density
Every episode is a time-synchronised, schema-typed recording you can replay topic by topic — vision, action and force aligned to one episode clock.
Native sample rates, as recorded
Rates are as recorded on production lines, not resampled.
400+ recorded topic types
Head, chest, torso, wrist and gripper cameras.
Joint position, velocity, current and command.
6-axis wrist F/T and end-effector wrench.
TF tree, camera intrinsics and extrinsics, SOP.
Every task breaks down into an ordered sequence of steps, and every step carries skill verbs from a controlled 140-term taxonomy.
Unitree G1, bimanual, stereo and monocular 720p, BrainCo hands. 22 annotated steps and 13 skill verbs in total; first eight shown.
Wearable first-person recordings of real human work — the cheapest route to task diversity. 40 scenarios, 1,400 recorded hours, nine rigs in rotation.
Washstand cleaning, cucumber cutting, dish washing, after-meal tidying, wok stir-frying.
Boxing, weighing, gluing and sealing at a packing station; manual pallet-jack transport.
Guest-room housekeeping: linen changes, fixture wiping, amenity arrangement.
Restroom cleaning in stereo; a six-camera outdoor waste-station rig.
Compact chest or head monocular rig for long household sessions.
Chest-mounted stereo pair with dual-video output per session.
Wide-FOV fisheye headset for cluttered interior work.
Six-camera cart-mounted array for outdoor and multi-worker scenes.
Our inverse-dynamics model recovers 6-DoF end-effector trajectories and contact points from this footage, turning passive video into robot-executable action labels.
The EdgeWAM engine scores every trajectory against rigid-body physics in both directions before it enters a delivery batch.
Does this episode obey physics?
Trajectories are checked against rigid-body kinematics, collision boundaries, joint limits and acceleration and jerk envelopes. Sensor dropouts, tracking jumps and teleoperator noise are flagged and quarantined before training — not discovered three epochs into your run.
What action produced this frame?
Recovers 6-DoF end-effector trajectory distributions and contact points directly from passive egocentric video, converting unlabelled archives into structured, robot-executable actions at a fraction of teleop cost.
| Metric | Benchmark target | Validated against | Status |
|---|---|---|---|
| Kinematic trajectory accuracy | MAE ≤ 3.1 mm / ≤ 2.8° | OptiTrack optical ground truth | Verified |
| Physical anomaly detection | 94.2% precision / 91.8% recall | Kinematic-limit and jerk injection | Verified |
| Contact force estimation | MAE ≤ 1.2 N (0.5–10 N window) | 6-axis F/T load-cell fixtures | Verified |
| Downstream sample efficiency | 38% fewer training hours | Policy convergence vs. raw data | Verified |
Same architecture, same compute budget, same task suite. Only the training corpus changes.
Task success rate after convergence
over unfiltered raw video. Physics-verified tokens remove the non-physical demonstrations a policy would otherwise learn to imitate — the largest single source of brittle failure at deployment.
Delivery and licensing
Slice the robot catalog by platform, domain, scenario, skill or capture configuration — or take a pre-cut dataset off the shelf.
Example dataset card
250 tasks across 90 home and retail scenarios — cleaning, tidying, clothing handling, refrigerator zoning, basket shopping and shelf restocking. Six capture configurations span Dex1 and BrainCo end effectors with monocular and stereo 720p vision.
Largest scenarios
Standardised capture in MCAP and LeRobot v3.0, with device serial, capture config and camera calibration bound to every record.
Automated GDPR/CCPA facial and biometric scrubbing, with signed commercial waivers from every operator and location.
EdgeWAM physics scoring, duplicate detection and a causal metadata graph for every episode.
S3/GCS transfer or physical shipment under a tier-1 SLA. You pay only for frames that clear the agreed thresholds.
Monthly robot-hours of net-new capture
Today's run-rate by line
Tell us the embodiment, the skills or the model capability you're training. We'll cut a free evaluation slice against your task list — full annotation, no commitment.
