Open model · Apache-2.0
RightWayUp
Full-circle image roll estimation with calibrated abstention, by ORTUS AI.
Find the right way up, or recognise when the image doesn’t define one.
What it is
RightWayUp is a neural network model, created by the ORTUS AI team, that estimates how far an image is rotated from upright, over the full 360° at 1° resolution. It returns the angle, a confidence score and an abstain flag, and it can write a corrected copy of the image.
We built it for our CHEQIT camera-health work. A CCTV camera can be online and recording but still useless, e.g. rolled 15° after a storm or hanging upside down after a maintenance visit, and uptime checks won’t notice. We needed a model that reads a single frame of dim, compressed footage on an ordinary server and says when it isn’t sure. We couldn’t find an open model that did all of that, so we built one.
It turned out strong on everyday photos as well, so we are releasing it with code and weights under Apache-2.0. Besides checking CCTV, dashcam and body-worn cameras for roll, it can straighten photo libraries and scans, or help with document capture.
- Angle
- 0 to 359° at 1° resolution, as a clockwise rotation you can undo.
- Confidence
- The share of the model’s probability that falls within ±10° of its answer.
- Abstain
- Raised below a threshold fixed on calibration data. Bare walls, sky and straight-down aerial views have no defined “up”, and the model says so.
- Full circle
- Sideways and upside-down images included. Camera-calibration models such as GeoCalib are built for about ±45° of roll.
It comes in six tiers, from Pico to Max. Every tier ships as ONNX models (FP32, FP16 and INT8) for Linux, Windows and macOS, and as Core ML packages for Apple silicon. Nano, Fast and Pico also come as batch-capable ONNX files, and Pico has a build that runs in a web browser, on the device, so nothing is uploaded; a browser demo is coming. Each file format has its own abstention thresholds, fixed on calibration data. There is a Python package and CLI, and the ONNX models run in any ONNX Runtime language binding.
BeforeTurned 180°, measured +179.1°

AfterCorrected to upright

Results
We froze the weights, tiers and thresholds first, on 29 September 2026. Then we opened a sealed final holdout, never used for training or for choosing models, and scored it once, with identical pixels for every model. It has 4,706 new photos (2,201 from COCO and 2,505 from Open Images), each also scored after simulated CCTV degradation, and real traffic and thermal cameras: the R-LiViT traffic camera in RGB and thermal (74 frames each) and the Aalborg thermal camera (150 frames). We compare with Woehrer 2026, the published leader on the COCO rotation benchmark (more on that benchmark below), on the same images. Accuracy is the share of images whose estimate is within 10° of the true angle, with every image answered.
On the 4,706 new photos, Max is within 10° on 93.0% of them, against 88.4% for Woehrer 2026, and gives 33 upside-down answers (off by 150° or more) against 185. With simulated CCTV degradation, Max scores 88.2% against 49.3%, and every tier in the sealed holdout (Nano to Max) stays at or above 79.6%. On the two unseen thermal cameras (224 images, a small sample), Max scores 87.5% against 50.9%. On RotBench, an independent benchmark, Max and Pro get every image right at all four rotations, which on RotBench-Small (50 photos) is above the published human baseline of 0.97 to 0.99 (more below).

Six tiers
The tiers trade speed for accuracy. Max is the accurate tier: a large model on every image, and on most hardware slower than Woehrer 2026 for a single image. Pro and Balanced are cascades: Fast answers first and passes its least-confident images to Max. Nano and Fast are the small, fast tiers. On clean photos they trail Woehrer 2026 or at best match it, and once the same photos carry simulated CCTV degradation they lead it by a wide margin. Pico, the smallest, also runs in a web browser. We chose it after the sealed holdout was opened, so it has development results only.
Sealed final holdout. Share of images within 10° of the true angle, all images answered. CPU latency is a single image (batch 1) on an Intel Core i7-1260P laptop, ONNX Runtime INT8, 4 threads; Balanced and Pro are derived from measured stage timings at their calibration routing share. Woehrer 2026 runs FP32 on the same CPU.
| Tier | Use it for | New COCO photos, within 10° (n = 2,201) | New Open Images photos, within 10° (n = 2,505) | Same photos with simulated CCTV degradation, COCO / Open Images | CPU, single image (batch 1) |
|---|---|---|---|---|---|
| MaxViT-L/14 at 280 px, every image | Best accuracy; photos, thermal and hard frames | 95.6% | 90.7% | 92.5% / 84.4% | 326 ms |
| ProFast → Max cascade, ~20% routed | Near-Max accuracy at lower cost | 95.5% | 90.5% | 92.4% / 84.1% | 94.8 ms |
| BalancedSame cascade, ~10% routed | Good accuracy at a fraction of Max’s cost | 94.0% | 89.1% | 91.6% / 83.2% | 62.2 ms |
| FastViT-S/14 at 224 px | CCTV, scenes, batch jobs | 90.4% | 84.5% | 89.3% / 78.9% | 29.7 ms |
| NanoViT-S/14 at 112 px | Edge devices; CCTV and scenes only | 87.3% | 75.5% | 86.5% / 73.5% | 7.73 ms |
| PicoViT-S/14 at 70 px | Web browsers and the smallest devices | Not part of the sealed holdout (chosen afterwards); see the development tests below | 3.21 ms | ||
| Woehrer 2026MambaOut-base CGD | Reference | 93.6% | 83.8% | 54.7% / 44.5% | 161 msFP32 |
Nano is built for scenes. On close-up or top-down object photos it is weaker (75.5% on the new Open Images photos, against 83.8% for Woehrer 2026), so use Balanced, Pro or Max for photo apps.
Sealed holdout, Nano to Max
Share of images within 10° of the true angle, all images answered. Upside-down errors (150° or more off) in brackets for Woehrer 2026 and Max. The max-area crop row shows the same COCO photos at the same angles, cropped the way the COCO rotation benchmark crops them.
| Test set | Woehrer 2026 | Max | Pro | Balanced | Fast | Nano |
|---|---|---|---|---|---|---|
| New COCO photos (n = 2,201) | 93.6% (40) | 95.6% (4) | 95.5% | 94.0% | 90.4% | 87.3% |
| Same photos, max-area crop | 94.1% (36) | 95.3% (5) | 95.0% | 93.2% | 90.2% | 88.0% |
| Same photos, simulated CCTV degradation | 54.7% (78) | 92.5% (9) | 92.4% | 91.6% | 89.3% | 86.5% |
| New Open Images photos (n = 2,505) | 83.8% (145) | 90.7% (29) | 90.5% | 89.1% | 84.5% | 75.5% |
| Same photos, simulated CCTV degradation | 44.5% (173) | 84.4% (24) | 84.1% | 83.2% | 78.9% | 73.5% |
| R-LiViT traffic camera, RGB (n = 74) | 98.6% (1) | 97.3% (0) | 93.2% | 87.8% | 87.8% | 77.0% |
| R-LiViT traffic camera, thermal (n = 74) | 37.8% (33) | 95.9% (0) | 95.9% | 95.9% | 90.5% | 74.3% |
| Aalborg thermal camera (n = 150) | 57.3% (37) | 83.3% (25) | 78.0% | 68.7% | 61.3% | 48.7% |
Max beats Woehrer 2026 significantly on six of the eight sets (95% paired bootstrap intervals) and ties on the other two: the max-area crop and the RGB traffic camera (74 images), where the smaller tiers trail Woehrer 2026. The paired intervals for each of these tiers are in the results on GitHub.
Held-out CCTV cameras
We also held out two test sets on 24 September, to score our earlier pre-release models, and scored them once more with the released models. This is their second use: the released models never trained on them and were never chosen with them. The first set is four unseen MEVA CCTV cameras, one of them thermal (380 views). The second combines DIODE indoor and outdoor scans, a MEVA CCTV camera and Poly Haven 3D renders (2,854 views). Each view appears once clean and once with simulated CCTV degradation. Two further MEVA cameras from the original set are left out, because our development tests used other clips from them.
Held-out sets, second use. Share of views within 10° of the true angle, all views answered; upside-down errors (150° or more off) in brackets.
| Test set | Woehrer 2026 | Deep-OAD | RightWayUp Max |
|---|---|---|---|
| Four unseen CCTV cameras, one thermal (n = 380 views) | 75.0% (47) | 82.4% (26) | 100.0% (0) |
| DIODE scans, a MEVA CCTV camera and Poly Haven renders (n = 2,854 views) | 70.6% (98) | 82.8% (113) | 98.8% (5) |
On both sets Max’s lead over Woehrer 2026 and over Deep-OAD, another open full-circle model, is significant (95% paired bootstrap intervals). We compare with Deep-OAD on these held-out sets only.
RotBench
RotBench is an independent benchmark that turns photos by 0°, 90°, 180° or 270° and publishes a human baseline. We scored it once on 30 September with the released models, after pre-registering the test, using its official protocol. Max and Pro get every image right on RotBench-Small (50 photos) and RotBench-Large (300 photos), with no upside-down answers. On RotBench-Small (50 photos) that is above the published human baseline (0.99, 0.99, 0.99 and 0.97 at the four rotations), while GPT-5, Gemini 2.5 Pro and o3, as published, score 0.40 to 0.81 on the turned photos.
Share of photos whose rotation is identified correctly (4-way accuracy), every photo answered; for RotBench-Large, the mean over the four rotations. The human and vision-language-model rows are as published by RotBench, not re-run by us. RotBench-Small has only 50 photos.
| Model | RotBench-Small (50 photos) | RotBench-Large (300 photos), mean | |||
|---|---|---|---|---|---|
| 0° | 90° | 180° | 270° | ||
| RightWayUp Max | 1.00 | 1.00 | 1.00 | 1.00 | 1.00 |
| RightWayUp Pro | 1.00 | 1.00 | 1.00 | 1.00 | 1.00 |
| RightWayUp Fast | 0.98 | 0.98 | 0.96 | 0.98 | 0.99 |
| RightWayUp Nano | 0.96 | 0.98 | 0.96 | 0.94 | 0.98 |
| Woehrer 2026 | 0.90 | 0.92 | 0.88 | 0.88 | 0.97 |
| Humansas published | 0.99 | 0.99 | 0.99 | 0.97 | – |
| GPT-5as published | 1.00 | 0.41 | 0.81 | 0.59 | – |
| Gemini 2.5 Proas published | 1.00 | 0.50 | 0.72 | 0.40 | – |
| o3as published | 1.00 | 0.45 | 0.70 | 0.48 | – |
Development tests
Development tests: never trained on, but used to choose the models, so they are less independent than the sealed and held-out sets. The MEVA row is a separate development panel of CCTV cameras, not the held-out cameras above. Share of views within 10° of the true angle, all views answered.
| Test | Woehrer 2026 | Max | Pro | Balanced | Fast | Nano | Pico |
|---|---|---|---|---|---|---|---|
| COCO rotation benchmark, our rebuild, five-seed mean (n = 5,150 views) | 98.0% | 98.8% | 98.7% | 97.7% | 96.5% | 96.0% | 92.2% |
| Same views, each saved once as JPEG at quality 90 | 30.2% | 98.4% | 98.2% | 97.5% | 96.4% | 95.9% | 92.2% |
| MEVA development cameras, never trained on (n = 646 views) | 77.5% | 99.8% | 99.8% | 99.7% | 98.8% | 98.0% | 96.4% |
In another development test, on photos re-rendered the way a physically rolled camera would see them, Max scores 83.6% against 77.6% for Woehrer 2026.
The COCO rotation benchmark
Woehrer 2026 (MambaOut-base CGD by Maximilian Woehrer, arXiv:2603.25351, MIT licence) is the published state of the art on the COCO rotation benchmark: 1,030 COCO photos, each rotated to a random angle under five test seeds. Its open release of code, weights and test list made these comparisons possible. On our rebuild of the benchmark, Max scores 98.8% against 98.0% for Woehrer 2026. Both are five-seed means, the figure the paper reports.
Two properties of the benchmark matter when its numbers are read as a guide to real cameras. Both come from a widely used protocol and aren’t specific to that paper.
The crop depends on the angle. The benchmark crops each rotated photo to the largest upright rectangle that fits inside it, as a lot of RotNet-style code does. The size and field of view of that rectangle change with the angle, and Woehrer 2026 trains on the same crop, while a rolled camera never provides that cue. We scored our sealed COCO photos both ways, at the same angles. Woehrer 2026 scores 93.6% with a fixed crop and 94.1% with the benchmark’s crop; Max scores 95.6% and 95.3%. Max’s lead is significant with the fixed crop and becomes a tie with the benchmark’s crop. The effect is small, and we haven’t proven how the crop helps.
JPEG. The benchmark rotates JPEG photos and keeps the rotated views without compressing them again, so each photo’s 8 × 8-pixel JPEG block pattern turns with the picture. A camera compresses its own frame, so its block pattern lines up with the frame edges. We saved every benchmark view once as an ordinary JPEG at quality 90, as a camera or phone would. Woehrer 2026 then falls from 98.0% to 30.2% (five-seed means), while RightWayUp Nano to Max score 95.9 to 98.4%. Our hypothesis is that Woehrer 2026 reads the rotated block pattern to find the exact angle; we haven’t proven that. Woehrer’s repository already notes that JPEG re-compression lowers accuracy; we measured by how much.
We report the benchmark both ways. The benchmark notes, the rebuild script and every test-set definition are on GitHub, and we welcome corrections.
Simulated CCTV degradation
CCTV frames are compressed, dim, noisy and blurred, and often infrared or low-resolution. CHEQIT watches cameras in exactly those conditions, so we trained on simulated versions of them and test the same way. Each new photo in the sealed holdout appears twice at the same angle, once clean and once passed through our simulated CCTV degradation (resolution loss, blur, infrared or greyscale, exposure, noise and JPEG compression). The degraded views are simulated, so we did not record damaged cameras for them.

Sealed final holdout. Share within 10° of the true angle, clean → simulated CCTV degradation, same photos and angles in both halves.
| Test set | Woehrer 2026 | RightWayUp Nano | RightWayUp Fast | RightWayUp Max |
|---|---|---|---|---|
| New COCO photos (n = 2,201 + 2,201 views) | 93.6% → 54.7% | 87.3% → 86.5% | 90.4% → 89.3% | 95.6% → 92.5% |
| New Open Images photos (n = 2,505 + 2,505 views) | 83.8% → 44.5% | 75.5% → 73.5% | 84.5% → 78.9% | 90.7% → 84.4% |
Woehrer 2026 was trained on clean photos, so this gap mostly reflects what each model was built for.
How it works

The backbone is a DINOv2 vision transformer (Apache-2.0, from Meta), fully fine-tuned for roll. It splits the frame into 14 × 14-pixel patches and reads each one as a token. The head outputs a probability for each of 360 one-degree bins. It is trained against a circular Gaussian target, so 359° and 1° count as neighbours. The peak is the angle, and the confidence is the probability within ±10° of the peak.
Training rotates images at random over the full 0 to 360° and crops them in a way that doesn’t depend on the angle, so the outline never gives the answer away. It also applies the simulated CCTV degradations described above. The released models were retrained so they don’t rely on the faint JPEG-grid and resampling traces that rotating a photo can leave, and every shipped file passes that check in its own runtime (details in the technical report).
Pico, Nano and Fast are single ViT-S/14 models. Nano reads images at 112 px, and Pico is Nano fine-tuned to read them at 70 px. Fast reads them at 224 px and was distilled from Max. Max is a ViT-L/14 that reads every image at 280 px. Balanced and Pro are cascades: Fast answers first, and its least-confident images go to Max (about 10% of calibration images for Balanced and 20% for Pro). If the answer is still uncertain, the tier abstains. Routing and abstention thresholds were fixed on calibration data and are not tuned per deployment. The routed share depends on the workload; hard images such as thermal frames are routed more often, which makes the cascades slower there.
Knowing when not to answer
Some images have no defined “up” (e.g. open sky, sand, a bare wall or a straight-down aerial view), so below a confidence threshold the model abstains rather than guess. We fixed the thresholds on separate calibration data, where the standard point answers about 90% of images, then applied them unchanged to every test. A strict point allows at most 1% wrong on calibration data, and abstention can be switched off. Each file format (ONNX FP32, FP16 and INT8, and Core ML) has its own thresholds, fitted on the same calibration data.
“I think abstention matters more than the headline number, because for a camera check a confident wrong answer costs more than a skipped frame.”
Vladimir Ilyash, CTO and co-founder, ORTUS AI

Sealed final holdout, standard operating point. Each cell: share of images the tier answers / share wrong (more than 10° off) among the answered images.
| Test set | Max | Pro | Balanced | Fast | Nano |
|---|---|---|---|---|---|
| New COCO photos (n = 2,201) | 91% / 2.1% | 91% / 2.2% | 89% / 3.4% | 89% / 4.6% | 89% / 7.1% |
| Same photos, simulated CCTV degradation | 86% / 3.0% | 86% / 3.1% | 87% / 4.0% | 87% / 4.5% | 88% / 6.8% |
| New Open Images photos (n = 2,505) | 74% / 1.2% | 74% / 1.7% | 75% / 3.2% | 73% / 4.4% | 71% / 10.3% |
| Same photos, simulated CCTV degradation | 64% / 2.4% | 64% / 2.8% | 66% / 4.1% | 67% / 5.5% | 70% / 11.1% |
| Aalborg thermal camera (n = 150) | 80% / 12.5% | 84% / 19.8% | 83% / 27.2% | 74% / 31.5% | 77% / 50.0% |
On unfamiliar photos the models answer less often rather than guess: on the new Open Images photos Max answers 74% of images and is wrong on 1.2% of those. Nano’s confidence is less dependable on photos (10.3% wrong among the Open Images photos it answers). Thermal cameras are the weak spot. On the unseen Aalborg thermal camera every tier in the table stays confident while wrong too often, and the strict point does not fix it. On thermal footage, use Max and don’t rely on the confidence alone.
Speed
We report two measurements separately. Single image (batch 1) is the median latency for one image processed on its own, which is what a per-camera check sees. Batched (batch N) is the sustained throughput when many images are processed together.
Pico, Nano and Fast take under 1 ms per image on a GPU: 0.55 ms, 0.55 ms and 0.68 ms for a single image (batch 1) with their one-image files on an NVIDIA RTX PRO 4500, TensorRT FP16. On an Apple M4 with Core ML FP16 on the Neural Engine they take 0.86 ms, 1.04 ms and 3.44 ms. On a laptop CPU (Intel Core i7-1260P, ONNX Runtime INT8, 4 threads) they take 3.21 ms, 7.73 ms and 29.7 ms, against 161 ms for Woehrer 2026 in FP32 on the same CPU. That makes Fast 5.4× and Pico 50× faster, INT8 against FP32.
Max is the accurate tier, not the fast one. For a single image it is slower than Woehrer 2026: 1.1 to 1.3× on the data-centre and workstation GPUs we measured and 1.6 to 5× on CPUs. Batched on a GPU, it gets through more images per second than Woehrer 2026, whose released graph takes one image at a time.

Single image (batch 1), median latency. GPUs run the batch-capable files; the one-image files of Pico, Nano and Fast are faster still (above). Balanced and Pro are derived from measured stage timings at the calibration routing share (10% / 20%); a workload that routes more images will be slower. Some rows are scaled to our earlier measurements on the same device; the method and every device we measured are in the benchmark results on GitHub.
| Hardware and runtime | Pico | Nano | Fast | Balanced (derived) | Pro (derived) | Max | Woehrer 2026 |
|---|---|---|---|---|---|---|---|
| NVIDIA RTX PRO 4500 GPU, TensorRT FP16 | 0.75 ms | 0.95 ms | 1.25 ms | 1.77 ms | 2.29 ms | 5.19 ms | 4.04 msONNX Runtime CUDA FP32 |
| NVIDIA RTX 5090 GPU, TensorRT FP16 | 0.68 ms | 0.86 ms | 0.98 ms | 1.48 ms | 1.98 ms | 4.99 ms | 3.73 msONNX Runtime CUDA FP32 |
| NVIDIA L4 GPU, TensorRT FP16 | 0.88 ms | 1.62 ms | 1.26 ms | 2.18 ms | 3.09 ms | 9.15 ms | 7.28 msONNX Runtime CUDA FP32 |
| Apple M4, Core ML FP16 on the Neural Engine | 0.86 ms | 1.04 ms | 3.44 ms | 8.78 ms | 14.1 ms | 53.5 ms | 138 msCPU FP32, 4 threads |
| Intel Core i7-1260P laptop CPU, ONNX Runtime INT8, 4 threads | 3.21 ms | 7.73 ms | 29.7 ms | 62.2 ms | 94.8 ms | 326 ms | 161 msFP32 |
| Intel Xeon Gold 6530 server CPU, ONNX Runtime INT8, 4 threads | 7.43 ms | 11.5 ms | 24.3 ms | 44.2 ms | 64.1 ms | 199 ms | 123 msFP32 |
| AMD EPYC 7663 server CPU, ONNX Runtime INT8, 4 threads | 12.9 ms | 18.6 ms | 42.4 ms | 95.8 ms | 149 ms | 534 ms | 150 msFP32 |
Batched (batch 64), sustained throughput in images per second, TensorRT FP16. Balanced and Pro are derived as above. Woehrer 2026’s released graph has a fixed batch of 1, so it has no batched figure; its single-image latency is in the table above.
| GPU | Pico | Nano | Fast | Balanced (derived) | Pro (derived) | Max |
|---|---|---|---|---|---|---|
| NVIDIA RTX PRO 4500 | 55,950 img/s | 25,737 img/s | 6,483 img/s | 2,631 img/s | 1,650 img/s | 443 img/s |
| NVIDIA RTX 5090 | 62,071 img/s | 34,706 img/s | 10,013 img/s | 4,827 img/s | 3,180 img/s | 932 img/s |
| NVIDIA L4 | 31,845 img/s | 12,889 img/s | 2,381 img/s | 947 img/s | 591 img/s | 157 img/s |
Batching amortises fixed per-call overhead and fills the GPU, so batched throughput is far higher than 1000 ÷ single-image latency. Every file format has its own abstention thresholds, fitted on the same calibration data. On CPU, INT8 is the recommended path.
Where it’s weaker
- In straight-down (nadir) aerial images, “up” points out of the picture, so there is no “up” within the two-dimensional image. The model should abstain. Don’t use it to orient nadir drone footage.
- Thermal cameras are the weakest case. On the unseen Aalborg thermal camera Max scores 83.3%, Nano and Fast do no better than Woehrer 2026, and the confidence of every tier we scored there is less reliable, so don’t rely on abstention for thermal footage. Most likely that is because thermal frames are a small part of the training data; fine-tuning on thermal footage should close much of the gap.
- Text-heavy images and people lying down are among the hardest cases, with occasional 180° flips.
- Nano and Fast trail Woehrer 2026 on clean COCO photos (87.3% and 90.4% against 93.6%), and Nano also on clean Open Images photos. They are built for scenes such as CCTV, rooms and streets, not close-up or top-down object photos. Pico gives up more accuracy for size and speed, and has development results only.
- If your images are never far from level and you need full camera calibration, GeoCalib is the better choice.
- It isn’t an IMU replacement, and we haven’t validated it on robots or vehicles.
- Training photos are mostly Flickr consumer photography, and the CCTV training data comes from one US campus dataset.
Data provenance
Rotation labels look free, since you can rotate any upright image, but you still need to know that the image was upright to begin with and that you have the rights to use it. Max, Fast and Pico were trained on about 1.24 million distinct images (1,244,222 in 1,245,604 training rows), each with a recorded source and licence, from sources that permit commercial use: 1.22 million photos plus DIODE scans, MEVA CCTV frames and Poly Haven 3D renders. Nano’s final training came before the Open Images V7 and CommonCatalog expansion and used 705,075 training rows. For the 629,551 COCO, Open Images and CommonCatalog photos, we re-checked each photo’s current Flickr licence; PASS photos carry the licence record of the PASS dataset. No customer data was used.
- PASS: Flickr photos without people; CC BY 4.0 dataset grant, with per-image CC BY 2.0 attribution.
- COCO: CC BY photos and a small public-domain-like set, with the current Flickr licence re-checked per image.
- Open Images (test subset and V7 train): CC BY photos and a small public-domain-like set, with the current Flickr licence re-checked per image, kept only where the recorded display orientation is upright and there is no EXIF rotation flag.
- CommonCatalog: YFCC100M photos whose current Flickr licence is exactly CC BY 2.0.
- DIODE: MIT.
- MEVA: CC BY 4.0.
- Poly Haven: CC0 assets, rendered by us. Our renders are published as the RightWayUp training renders dataset (CC BY 4.0).
We don’t redistribute the photos and frames, only our own renders. The published manifests trace every image to its source, with author credit where the licence requires it, so anyone can rebuild the set under the original terms. Near-duplicates of test and benchmark photos were excluded from training. The DINOv2 backbone is Apache-2.0. Meta doesn’t warrant the rights trail of its pretraining data, and our data card says so.
Use it
pip install rightwayup
rightwayup predict frame.jpg # angle_cw, confidence, abstain
rightwayup fix photo.jpg # writes an upright copy; leaves it alone if it abstainsfrom PIL import Image
from rightwayup import Orienter
o = Orienter(tier="max") # downloads ONNX weights from the Hugging Face Hub
r = o.predict("frame.jpg") # r.angle_cw, r.confidence, r.abstain, r.routed
img = Image.open("photo.jpg")
upright = o.correct(img, snap=90) # rotated back to upright, snapped to the nearest 90°angle_cw is the clockwise rotation of the image content, so rotating by that amount counter-clockwise restores it. Options for the tier (pico, nano, fast, balanced, pro or max), the abstention point (standard, strict or off), the device and batching are in the README on GitHub.
Licence and citation
Code and weights are released under Apache-2.0, including the Max tier, and commercial use is allowed. Keep the LICENSE and NOTICE files when redistributing. “ORTUS AI” and “RightWayUp” are trademarks and are not covered by the licence; please give modified or retrained models a different name.
If RightWayUp helps your product, research or project, we would appreciate a mention such as “Orientation by RightWayUp from ORTUS AI” with a link to this page, or a citation. This request adds no condition to the licence.
@misc{ortusai2026rightwayup,
title = {RightWayUp: Full-circle image roll estimation with calibrated abstention},
author = {ORTUS AI},
year = {2026},
url = {https://github.com/ortusaitech/rightwayup}
}Links
- GitHub github.com/ortusaitech/rightwayupCode, examples, benchmark scripts, the technical report, the benchmark notes and full result tables.
- Hugging Face huggingface.co/ortusai/rightwayupModel weights for all six tiers, the model card and instructions for running the ONNX files directly.
- PyPI pypi.org/project/rightwayupPython package and CLI.
Built for CHEQIT
We built RightWayUp for our CHEQIT camera-health work, to spot from a single frame that a camera has been knocked, rolled or remounted, or was mounted at the wrong angle in the first place, and to say by how many degrees. That saves maintenance visits and guesswork. CHEQIT is ORTUS AI’s on-prem monitoring for CCTV fleets.

If it fails on your images, we’d like to hear about it at hello@ortusai.io. If you need it adapted to thermal, fisheye or document imagery, or tuned for specific edge hardware, please reach out. We’re happy to work with you on it.
Credits
- Roller-coaster footage in the video: “Huracan POV (Front Row)” and “Impulse POV” by Canobie Coaster, CC BY 3.0, via Wikimedia Commons (Huracan, Impulse). Modified: levelled by RightWayUp.
- Action footage in the video: Pexels (Pexels License), Alex Moliski (mountain bike) and other Pexels creators (FPV and alpine coaster).
- Scenes in the video and the stills on this page: Poly Haven HDRIs (CC0). Training-domain tile in the video: MEVA (Kitware / IARPA, CC BY 4.0).
- Music and sound effects: generated with ElevenLabs on a paid plan that allows commercial use. ElevenLabs says its music model is trained on licensed stems and music, built in partnership with artists, labels and publishers. Font: Inter (SIL Open Font License).
- Comparison models, each scored from its public release: Woehrer 2026 (code, paper, MIT licence); Deep-OAD, on the held-out sets only (code, scored via its public ONNX export). Also mentioned: GeoCalib (code).
All figures come from the RightWayUp 1.0 release weights. The sealed final holdout was opened once, after the weights, tiers and thresholds were frozen on 29 September 2026, and scored once; the development tests, which were also used to choose the models, were scored with the same weights. The held-out CCTV sets were used once before, on 24 September, to score our earlier pre-release models, and scored once more with the released models. Pico was chosen after the holdout was opened, so it has development results only. “Within 10°” is the share of images whose estimate is within 10° of the true angle.