Skip to content
ucr.

Proof

What we have measured, and how.

Synthetic data is only as good as the simulator behind it. Here is every check we have run, the method behind it, and what it does not prove yet.

Verified · Jun 2026

Flies like a real FPV drone

Same outputs as Betaflight firmware
What we did

Fed identical stick and gyro inputs to our controller and to the real Betaflight code, then compared their outputs step by step.

Result

The P and I terms agree to about 1e-4. The drone in our sim is controlled the way a real FPV drone is.

Not yet proven

This checks the control code, not the motors and props. The next result covers those.

Verified · Jun 2026

More accurate than a published physics model

Lower error on all 4 force and torque axes
What we did

Replayed real quadrotor flights from the NeuroBEM dataset through our physics and measured the force and torque error.

Result

Lower error than the paper's blade-element physics model on all four force and torque axes.

Not yet proven

The shape of roll and pitch torque does not match closely yet, and we have not compared against a drone we own.

Verified · Jul 2026

Already scores real AI pilots

Hover held within 6.2 cm
What we did

Loaded SimpleFlight's pretrained flight policy, flew it through our simulator over our Python interface, and measured its error against ground truth.

Result

Hover held within 6.2 cm. A figure-8 flown within 8.0 cm. This is our evaluation loop working end to end.

Not yet proven

Above about 2.5 m/s our sim drifts from theirs, and we have not yet compared against their published numbers.

In progress · due Jan 2027

Does a detector trained on our data find real drones?

Due Jan 2027
What we are doing

Training a standard detector on real photos only, then on real plus our synthetic data, and testing both on real drone photos from the public DUT Anti-UAV dataset.

What we will publish

The method, the code, the dataset description and the numbers over several runs, whether they help us or not.

Status

Step 1 of 3 done: the real-only baseline is reproduced at 0.90 to 0.91 mAP@50 over two runs, against a published 0.915. Next we add our synthetic data and measure the change. Get the result by email →

Our testing ladder

Four levels of trust, climbed in order.

  1. Done for physics

    1. Does it move right?

    Flight controller and physics compared with firmware and real flight logs.

  2. Started

    2. Do sensors look real?

    Simulated gyro noise and drift checked against the statistics of real chips.

  3. In progress

    3. Does it transfer?

    Models trained on our data, tested on real imagery. This is the number that matters most.

  4. Planned

    4. Does it fly?

    Models tested on real hardware connected to the simulator, then in real flight.

Our rules

How we keep this honest.

No claim without a number

If we haven't measured it, the site says "planned".

Same settings, same output

Every run is seeded. Automated tests check that the same scenario always produces the same flight.

Method first, then results

Public benchmarks, published protocols and code, so anyone can check the result.