Trekion
Early access

Simulation infrastructure for physical AI

Trekion rebuilds real facilities as photoreal, physics-accurate simulation. Train and test your robots there before they reach the floor: environments and assets, evaluation, and benchmarking on one pipeline.

policy-health

Current success rate

86%

Across all validation test cases

Latest checkpoint

Model · Epoch 600

Best success rate

100% pass

Validation set5/5 passed100%
Aisle congestion5/5 passed100%
Floor friction5/5 passed100%
Sensor delay4/5 passed80%

from the evaluation workspace

Built on open standardsOpenUSDIsaac SimMuJoCoROSGazebo
The problem

Testing robots in the real world does not scale.

01

Robot data is never free

Language models trained on data that already existed on the internet. Robot data has to be collected each time: hardware hours, supervision, facility access, or payments to a data provider.

02

Generic sim cannot be trusted

Simulation is the obvious answer, but stock simulators carry too much sim-to-real gap: game-like rendering, approximate contact, nobody's actual site. A sim you cannot trust returns confident, wrong answers.

03

Accurate sim changes the economics

Close the gap and the economics change: evaluations you can rely on, run in minutes instead of hundreds of hardware hours, before the fleet is committed.

On hardware-testing cost, Genesis AI has published figures running to hundreds of hours per real-world evaluation pass.

The precedent

Self-driving already went through this.

Waymo has described driving on the order of 20 billion miles in simulation against tens of millions on real roads. No autonomy program has reached deployment on physical miles alone. And driving is the easy version: one body, one road network, one job. A warehouse floor holds many embodiments, many tasks, and none of driving's free data.

A week on hardware

≈ 40 rollouts

An afternoon in sim

10,000 rollouts

Failures get grouped automatically

Failed runs are clustered by cause and ranked by cost, so the pattern reaches you before the footage does.

Stalls in narrow aisle congestion214
Misses pallet edge under low light156
Overshoots dock alignment98
The evidence

Generic simulation misses the site.

In published work on forklift perception, a model trained only in a generic simulator managed 49.4% recall on photos from a real warehouse. World-model synthetic data lifted the same task to 84.7% recall. Grounding the data in the actual site, its floors, its labels, its lighting, took precision to 99.5%.

Generic simulation carries you partway. The distance that remains is the site itself, and building that half is exactly what our pipeline does.

Source: SoftServe and NVIDIA with Toyota Material Handling Europe.

perception-recall · forklift study
Generic simulator49.4% recall
World-model synthetic84.7% recall
Grounded in the site99.5% precision

forklift perception, real warehouse photos. same task, same model family. figures published by SoftServe, NVIDIA and Toyota Material Handling Europe.

From our pipeline

One captured scene, two views.

A room we captured, reconstructed, and loaded into Isaac Sim, rendering in real time at 60 FPS. On the left, the scene a camera would see. On the right, the same scene as the physics solver sees it. Most simulation gives you one or the other. A scene has to hold up as both to be worth testing in.

A reconstructed living room rendered photorealistically in Isaac Sim at 60 FPS
What the camera seesRTX real-time · 60 FPS
The same reconstructed room shown as collision geometry for the physics solver
What the solver seescollision geometry · PhysX
The shape of it

One pipeline. Three products.

Everything rests on one capability: turning a working facility into a simulation that behaves like it. The environments are what teams buy first. The evaluation is what proves the worlds are honest. The loop is what keeps them that way.

Every run, and eventually every deployment, feeds the pipeline. The worlds get harder to compete with each time they are used.

Working facility

scans, video, floor plans

Real-to-sim pipeline

capture · reconstruct · physics · randomise · report

Assets & environments

Evaluation

Benchmarking

field results return as scenarios
Products

Three products on one pipeline.

01Early access

Assets & environments

Environments built from real facilities.

Library and generation

A repository of physics-ready assets and environments, plus an LLM-driven interface that generates new ones from a prompt or a spec.

Deployment-specific real-to-sim

Your actual site reconstructed as a photoreal, physics-accurate scene for training and validation before commissioning.

Request access
02Early access

Evaluation

Where a go-live gets signed off.

Scenario-based evaluation

Your policies run against structured suites built from scenarios you define, with subgoal-level grading and failure clustering.

Deployment-environment evaluation

The same harness inside the twin of your specific site, so the report reads like a go-live decision for that floor.

Request access
03Early access

Benchmarking

See where your policy stands.

Against your own history

Every checkpoint measured on identical suites and seeds, so progress and regressions are visible version over version.

Open benchmark

Compare your policies against open-source and custom baselines on the same environments, on one scoreboard.

Request access
How it works

From a real building to a working simulation.

01

Capture

The floor is scanned as it runs: layout, racking, surfaces, lighting, and the operating patterns around them.

From the real building

Video, photos, and scans of the live site, not a CAD idealisation.

Hours, not months

Capture-to-scene turnaround measured in hours.

Minimal disruption

No shutdowns; the floor keeps working.

facility.usd
OpenUSDcollidersarticulationsemanticsmetric scale
Prim pathTypeCount
/World/facilityscene1
/World/facility/rackingmesh248
/World/facility/floormesh1
/World/facility/palletsmesh96
/World/agents/amr_01articulation1
export targetsIsaac Sim · MuJoCo · Gazebo
02

Reconstruct

Neural reconstruction turns the capture into a scene that is photoreal to a camera and solid to a physics solver at the same time. Most simulation gives you one or the other; a scene has to be both to be worth training in.

What the camera sees

Lighting, reflections, glare, and sensor noise a vision stack will meet.

What the solver sees

Collision geometry and articulation underneath the same scene.

Metric scale

Dimensions that survive contact, not just look right.

reconstruction · site-04

01Raw capture

video, photos, and scans of the working floor

02Neural reconstruction

splatting into photoreal geometry

03Mesh and colliders

geometry a physics solver can use

04Semantics and scale

racking, lanes, SKUs, metric dimensions

both views, one scenewhat the camera sees · what the solver sees
03

Physics-ready assets

Every object carries measured physical properties: mass, friction pairs, articulation, collision hulls. The SKUs your fleet will actually touch, not a stock library's approximations.

Measured, not defaulted

Inertial and contact properties set from the real thing.

Real inventory

Your pallets, totes, and cages, worn the way they are worn.

Portable

USD, MJCF, and URDF out of the same asset.

asset · eur-pallet-01
physics-readymetric scalereal SKU
Mass24.75 kg
Friction0.45 on floor · 0.48 on gripper
Articulationrigid body
Collision meshconvex decomposition, 214 hulls
Materialwood, worn
FormatsUSD · MJCF · URDF

every asset ships with measured physical properties, not defaults.

04

Randomise

The scene multiplies into thousands of structured variants: aisle widths, traffic, lighting, load states. The conditions that end pilots, generated on purpose, usable for training or for evals.

Structured sweeps

Parameterised suites, not random noise.

The long tail on demand

Rare conditions become repeatable test cases.

Deterministic

Same seed, same run, every time.

scenario-inspector

Aisle congestion

suite: aisle-congestion

Structured variations of aisle width, traffic density, and floor friction across the captured facility.

Environment parameters

Aisle width: 1.4 m – 2.6 mTraffic density: 0 – 12 agentsFloor friction: 0.4 – 0.9seed: 101

Environments (5)

IDAisle widthAgents
congestion-011.4 m12
congestion-021.8 m8
congestion-032.2 m4
05

Report

Not a score, a diagnosis. Tasks decompose into subgoals graded one by one, so the report names the exact step that broke and the state of the stack when it happened.

Subgoal grading

Depart, navigate, yield, align, dock, each scored separately.

Failure clusters

Grouped by cause, ranked by what they cost.

Replayable

Step through any episode frame by frame.

episode-142
Move pallet from aisle 7 to dock 2failed

Atomic subgoals

departnavigate aisleyield to agentalign to palletdock
#164INFOpath clear, entering aisle 7
#165WARNagent detected at 1.2 m, yield triggered
#166WARNsubgoal failed: yield window exceeded
#182INFOrecovered, resuming to dock 2
The loop

Every run improves the simulation.

Field results flow back in. A failure on the floor becomes a scenario in the suite. A success confirms the physics. Checkpoint after checkpoint gets measured against the same conditions, so regressions surface in simulation instead of on hardware.

This is what separates infrastructure from a services engagement: the worlds compound. Each customer run leaves the simulation more accurate than it found it.

checkpoint-sweep
Checkpoints side by side5 policies · 1,199 episodes
nav-3b-v281.7%
nav-2b-sft-v384.2%
nav-5b-v379.3%
nav-6b-v274.8%
nav-2b-sft-v266.9%

nav-2b-sft-v3 leads by 2.5 points. more post-training did not help.

Verticals

Built at facility scale.

Fleets, congestion, mixed human and robot traffic, and layouts that shift every quarter. Facility-scale problems, not benchtop ones. Research labs rebuild a tabletop; we rebuild the building.

Warehouse robotics · liveManufacturing · soonData centre robotics · soon
failure-suites

Coverage

Validation set100%
Aisle congestion100%
Floor friction100%
Sensor delay80%
Ambient lighting80%
Partial occlusion60%
Pallet placement80%
Scenario library
01

Navigate a narrow aisle

02

Dock to a pallet

03

Yield to a pedestrian

04

Handle a blocked lane

05

Traverse a shift change

06

Recover from lost localization

07

Cross a busy intersection

08

Pick from a mixed rack

Where we are

Early access.

The pipeline runs today: environments reconstructed from real capture, physics-ready assets, structured evaluation with detailed reports. A self-serve workspace and more verticals are on the way. Access is gated while we work with initial teams.

Real-to-sim from capturePhysics-ready assetsScenario suitesEvaluation reportsUSD · Isaac · MuJoCo · Gazebo
FAQ

Common questions.

What is Trekion?

Simulation infrastructure for physical AI. We reconstruct real facilities into photoreal, physics-accurate simulation, run policy evaluations and benchmarks inside them, and feed deployment data back so the worlds keep improving.

How is this different from benchmark evaluation?

Benchmark platforms score your policy on standard task suites. We score it inside the facility it will actually ship to: that site's racking, floors, lighting, and traffic. The output reads like a go-live decision, not a leaderboard entry.

Why does site-specific matter? Is generic sim not enough?

Published work on forklift perception found a simulator-only model reached 49.4% recall on real warehouse photos, world-model synthetic data lifted it to 84.7%, and grounding in the actual site pushed precision to 99.5%. Generic gets you partway. The site is the rest.

Which stacks does it work with?

Scenes and assets export to OpenUSD, Isaac Sim, MuJoCo, and Gazebo. Evaluation is policy-agnostic across VLA and navigation stacks through a standard harness.

Do you need access to my facility?

For a site twin, yes: scans, video, or floor plans. For generated environments, asset packs, and evaluation on our library, you can start without any facility access.

What stage is the product at?

All three products are in early access, gated while we work with initial teams. The pipeline itself, capture to evaluation report, runs today.

Get started

Tell us what you are building.

A 30-minute technical call. We will map the pipeline to your stack and put together a sample scene from your spec.

  • 30-minute technical call
  • Straight to the founders
  • A sample scene from your use case
Which products