Open model · Apache-2.0

RightWayUp

Full-circle image roll estimation with calibrated abstention, by ORTUS AI.

Find the right way up, or recognise when the image doesn’t define one.

GitHub Hugging Face PyPI
pip install rightwayup
RightWayUp in 1 minute 58 seconds, with music and no narration. Every RightWayUp angle and number on screen is real output of the release weights; Woehrer 2026 was scored on the same images, and the human and AI rows on the RotBench card are as published by RotBench. The CCTV knock is simulated, and the CHEQIT alert card is an illustration. Footage and music credits are at the end of this page.

What it is

RightWayUp is a neural network model, created by the ORTUS AI team, that estimates how far an image is rotated from upright, over the full 360° at 1° resolution. It returns the angle, a confidence score and an abstain flag, and it can write a corrected copy of the image.

We built it for our CHEQIT camera-health work. A CCTV camera can be online and recording but still useless, e.g. rolled 15° after a storm or hanging upside down after a maintenance visit, and uptime checks won’t notice. We needed a model that reads a single frame of dim, compressed footage on an ordinary server and says when it isn’t sure. We couldn’t find an open model that did all of that, so we built one.

It turned out strong on everyday photos as well, so we are releasing it with code and weights under Apache-2.0. Besides checking CCTV, dashcam and body-worn cameras for roll, it can straighten photo libraries and scans, or help with document capture.

Angle
0 to 359° at 1° resolution, as a clockwise rotation you can undo.
Confidence
The share of the model’s probability that falls within ±10° of its answer.
Abstain
Raised below a threshold fixed on calibration data. Bare walls, sky and straight-down aerial views have no defined “up”, and the model says so.
Full circle
Sideways and upside-down images included. Camera-calibration models such as GeoCalib are built for about ±45° of roll.

It comes in six tiers, from Pico to Max. Every tier ships as ONNX models (FP32, FP16 and INT8) for Linux, Windows and macOS, and as Core ML packages for Apple silicon. Nano, Fast and Pico also come as batch-capable ONNX files, and Pico has a build that runs in a web browser, on the device, so nothing is uploaded; a browser demo is coming. Each file format has its own abstention thresholds, fixed on calibration data. There is a Python package and CLI, and the ONNX models run in any ONNX Runtime language binding.

BeforeTurned 180°, measured +179.1°

An upside-down view of a flat-topped mountain above a meadow. RightWayUp Max reads a measured roll of plus 179.1 degrees, error 0.9 degrees, confidence 0.90, while its probability ring peaks at the bottom.

AfterCorrected to upright

The same view turned back upright by the model's estimate: meadow below, mountain and sky above, marked corrected to upright.
Lilienstein, Saxony. Poly Haven HDRI (CC0); we extracted a level view and turned it 180° clockwise. RightWayUp Max measured +179.1° (error 0.9°) with confidence 0.90, and the view was turned back by that estimate. The ring is the model’s actual output over 360°. Still from the launch video.

Results

We froze the weights, tiers and thresholds first, on 29 September 2026. Then we opened a sealed final holdout, never used for training or for choosing models, and scored it once, with identical pixels for every model. It has 4,706 new photos (2,201 from COCO and 2,505 from Open Images), each also scored after simulated CCTV degradation, and real traffic and thermal cameras: the R-LiViT traffic camera in RGB and thermal (74 frames each) and the Aalborg thermal camera (150 frames). We compare with Woehrer 2026, the published leader on the COCO rotation benchmark (more on that benchmark below), on the same images. Accuracy is the share of images whose estimate is within 10° of the true angle, with every image answered.

On the 4,706 new photos, Max is within 10° on 93.0% of them, against 88.4% for Woehrer 2026, and gives 33 upside-down answers (off by 150° or more) against 185. With simulated CCTV degradation, Max scores 88.2% against 49.3%, and every tier in the sealed holdout (Nano to Max) stays at or above 79.6%. On the two unseen thermal cameras (224 images, a small sample), Max scores 87.5% against 50.9%. On RotBench, an independent benchmark, Max and Pro get every image right at all four rotations, which on RotBench-Small (50 photos) is above the published human baseline of 0.97 to 0.99 (more below).

Summary card titled Why it’s different. 360°: full circle, sideways and upside down. 93.0%: new photos it had never seen, Max, Woehrer 2026 88.4%, sealed holdout. 88.2%: simulated CCTV degradation, Max, Woehrer 2026 49.3%, same photos. 0.68 ms: per image, Fast tier, TensorRT FP16, RTX PRO 4500, batch 1. Apache-2.0: open source, open weights, six tiers, Pico to Max. 33 vs 185: upside-down answers, Max against Woehrer 2026, errors of 150° or more, same photos. Footer: 100.0% within 10° on 4 CCTV cameras it had never seen, RightWayUp Max, Woehrer 2026 75.0%, 380 views, clean and simulated degradation, held-out cameras, test set used a second time.
Summary card from the launch video. The photo figures are the sealed final holdout (4,706 new COCO and Open Images photos), and the camera line is the held-out CCTV test; both are in the tables below.

Six tiers

The tiers trade speed for accuracy. Max is the accurate tier: a large model on every image, and on most hardware slower than Woehrer 2026 for a single image. Pro and Balanced are cascades: Fast answers first and passes its least-confident images to Max. Nano and Fast are the small, fast tiers. On clean photos they trail Woehrer 2026 or at best match it, and once the same photos carry simulated CCTV degradation they lead it by a wide margin. Pico, the smallest, also runs in a web browser. We chose it after the sealed holdout was opened, so it has development results only.

Sealed final holdout. Share of images within 10° of the true angle, all images answered. CPU latency is a single image (batch 1) on an Intel Core i7-1260P laptop, ONNX Runtime INT8, 4 threads; Balanced and Pro are derived from measured stage timings at their calibration routing share. Woehrer 2026 runs FP32 on the same CPU.

Tier Use it for New COCO photos, within 10° (n = 2,201) New Open Images photos, within 10° (n = 2,505) Same photos with simulated CCTV degradation, COCO / Open Images CPU, single image (batch 1)
MaxViT-L/14 at 280 px, every imageBest accuracy; photos, thermal and hard frames95.6%90.7%92.5% / 84.4%326 ms
ProFast → Max cascade, ~20% routedNear-Max accuracy at lower cost95.5%90.5%92.4% / 84.1%94.8 ms
BalancedSame cascade, ~10% routedGood accuracy at a fraction of Max’s cost94.0%89.1%91.6% / 83.2%62.2 ms
FastViT-S/14 at 224 pxCCTV, scenes, batch jobs90.4%84.5%89.3% / 78.9%29.7 ms
NanoViT-S/14 at 112 pxEdge devices; CCTV and scenes only87.3%75.5%86.5% / 73.5%7.73 ms
PicoViT-S/14 at 70 pxWeb browsers and the smallest devicesNot part of the sealed holdout (chosen afterwards); see the development tests below3.21 ms
Woehrer 2026MambaOut-base CGDReference93.6%83.8%54.7% / 44.5%161 msFP32

Nano is built for scenes. On close-up or top-down object photos it is weaker (75.5% on the new Open Images photos, against 83.8% for Woehrer 2026), so use Balanced, Pro or Max for photo apps.

Sealed holdout, Nano to Max

Share of images within 10° of the true angle, all images answered. Upside-down errors (150° or more off) in brackets for Woehrer 2026 and Max. The max-area crop row shows the same COCO photos at the same angles, cropped the way the COCO rotation benchmark crops them.

Test set Woehrer 2026 Max Pro Balanced Fast Nano
New COCO photos (n = 2,201)93.6% (40)95.6% (4)95.5%94.0%90.4%87.3%
Same photos, max-area crop94.1% (36)95.3% (5)95.0%93.2%90.2%88.0%
Same photos, simulated CCTV degradation54.7% (78)92.5% (9)92.4%91.6%89.3%86.5%
New Open Images photos (n = 2,505)83.8% (145)90.7% (29)90.5%89.1%84.5%75.5%
Same photos, simulated CCTV degradation44.5% (173)84.4% (24)84.1%83.2%78.9%73.5%
R-LiViT traffic camera, RGB (n = 74)98.6% (1)97.3% (0)93.2%87.8%87.8%77.0%
R-LiViT traffic camera, thermal (n = 74)37.8% (33)95.9% (0)95.9%95.9%90.5%74.3%
Aalborg thermal camera (n = 150)57.3% (37)83.3% (25)78.0%68.7%61.3%48.7%

Max beats Woehrer 2026 significantly on six of the eight sets (95% paired bootstrap intervals) and ties on the other two: the max-area crop and the RGB traffic camera (74 images), where the smaller tiers trail Woehrer 2026. The paired intervals for each of these tiers are in the results on GitHub.

Held-out CCTV cameras

We also held out two test sets on 24 September, to score our earlier pre-release models, and scored them once more with the released models. This is their second use: the released models never trained on them and were never chosen with them. The first set is four unseen MEVA CCTV cameras, one of them thermal (380 views). The second combines DIODE indoor and outdoor scans, a MEVA CCTV camera and Poly Haven 3D renders (2,854 views). Each view appears once clean and once with simulated CCTV degradation. Two further MEVA cameras from the original set are left out, because our development tests used other clips from them.

Held-out sets, second use. Share of views within 10° of the true angle, all views answered; upside-down errors (150° or more off) in brackets.

Test set Woehrer 2026 Deep-OAD RightWayUp Max
Four unseen CCTV cameras, one thermal (n = 380 views)75.0% (47)82.4% (26)100.0% (0)
DIODE scans, a MEVA CCTV camera and Poly Haven renders (n = 2,854 views)70.6% (98)82.8% (113)98.8% (5)

On both sets Max’s lead over Woehrer 2026 and over Deep-OAD, another open full-circle model, is significant (95% paired bootstrap intervals). We compare with Deep-OAD on these held-out sets only.

RotBench

RotBench is an independent benchmark that turns photos by 0°, 90°, 180° or 270° and publishes a human baseline. We scored it once on 30 September with the released models, after pre-registering the test, using its official protocol. Max and Pro get every image right on RotBench-Small (50 photos) and RotBench-Large (300 photos), with no upside-down answers. On RotBench-Small (50 photos) that is above the published human baseline (0.99, 0.99, 0.99 and 0.97 at the four rotations), while GPT-5, Gemini 2.5 Pro and o3, as published, score 0.40 to 0.81 on the turned photos.

Share of photos whose rotation is identified correctly (4-way accuracy), every photo answered; for RotBench-Large, the mean over the four rotations. The human and vision-language-model rows are as published by RotBench, not re-run by us. RotBench-Small has only 50 photos.

Model RotBench-Small (50 photos) RotBench-Large (300 photos), mean
0° 90° 180° 270°
RightWayUp Max1.001.001.001.001.00
RightWayUp Pro1.001.001.001.001.00
RightWayUp Fast0.980.980.960.980.99
RightWayUp Nano0.960.980.960.940.98
Woehrer 20260.900.920.880.880.97
Humansas published0.990.990.990.97–
GPT-5as published1.000.410.810.59–
Gemini 2.5 Proas published1.000.500.720.40–
o3as published1.000.450.700.48–

Development tests

Development tests: never trained on, but used to choose the models, so they are less independent than the sealed and held-out sets. The MEVA row is a separate development panel of CCTV cameras, not the held-out cameras above. Share of views within 10° of the true angle, all views answered.

Test Woehrer 2026 Max Pro Balanced Fast Nano Pico
COCO rotation benchmark, our rebuild, five-seed mean (n = 5,150 views)98.0%98.8%98.7%97.7%96.5%96.0%92.2%
Same views, each saved once as JPEG at quality 9030.2%98.4%98.2%97.5%96.4%95.9%92.2%
MEVA development cameras, never trained on (n = 646 views)77.5%99.8%99.8%99.7%98.8%98.0%96.4%

In another development test, on photos re-rendered the way a physically rolled camera would see them, Max scores 83.6% against 77.6% for Woehrer 2026.

The COCO rotation benchmark

Woehrer 2026 (MambaOut-base CGD by Maximilian Woehrer, arXiv:2603.25351, MIT licence) is the published state of the art on the COCO rotation benchmark: 1,030 COCO photos, each rotated to a random angle under five test seeds. Its open release of code, weights and test list made these comparisons possible. On our rebuild of the benchmark, Max scores 98.8% against 98.0% for Woehrer 2026. Both are five-seed means, the figure the paper reports.

Two properties of the benchmark matter when its numbers are read as a guide to real cameras. Both come from a widely used protocol and aren’t specific to that paper.

The crop depends on the angle. The benchmark crops each rotated photo to the largest upright rectangle that fits inside it, as a lot of RotNet-style code does. The size and field of view of that rectangle change with the angle, and Woehrer 2026 trains on the same crop, while a rolled camera never provides that cue. We scored our sealed COCO photos both ways, at the same angles. Woehrer 2026 scores 93.6% with a fixed crop and 94.1% with the benchmark’s crop; Max scores 95.6% and 95.3%. Max’s lead is significant with the fixed crop and becomes a tie with the benchmark’s crop. The effect is small, and we haven’t proven how the crop helps.

JPEG. The benchmark rotates JPEG photos and keeps the rotated views without compressing them again, so each photo’s 8 × 8-pixel JPEG block pattern turns with the picture. A camera compresses its own frame, so its block pattern lines up with the frame edges. We saved every benchmark view once as an ordinary JPEG at quality 90, as a camera or phone would. Woehrer 2026 then falls from 98.0% to 30.2% (five-seed means), while RightWayUp Nano to Max score 95.9 to 98.4%. Our hypothesis is that Woehrer 2026 reads the rotated block pattern to find the exact angle; we haven’t proven that. Woehrer’s repository already notes that JPEG re-compression lowers accuracy; we measured by how much.

We report the benchmark both ways. The benchmark notes, the rebuild script and every test-set definition are on GitHub, and we welcome corrections.

Simulated CCTV degradation

CCTV frames are compressed, dim, noisy and blurred, and often infrared or low-resolution. CHEQIT watches cameras in exactly those conditions, so we trained on simulated versions of them and test the same way. Each new photo in the sealed holdout appears twice at the same angle, once clean and once passed through our simulated CCTV degradation (resolution loss, blur, infrared or greyscale, exposure, noise and JPEG compression). The degraded views are simulated, so we did not record damaged cameras for them.

Six tiles of the same Venice palazzo, each with a different simulated degradation: low resolution, out of focus, night with sensor noise, infrared greyscale, heavy compression, and haze with low contrast. Every tile is marked upright with the correction RightWayUp Max applied. Headline: built for real camera conditions; RightWayUp Max on 4,706 new photos with simulated CCTV degradation, within 10°, sealed holdout, 88.2%, Woehrer 2026 49.3%.
Venice. Poly Haven HDRI (CC0); level view extracted, degradations and rotations applied by us; every correction is real RightWayUp Max output. The headline line is the sealed-holdout result on the 4,706 new photos with simulated CCTV degradation, shown per photo set in the table below.

Sealed final holdout. Share within 10° of the true angle, clean → simulated CCTV degradation, same photos and angles in both halves.

Test set Woehrer 2026 RightWayUp Nano RightWayUp Fast RightWayUp Max
New COCO photos (n = 2,201 + 2,201 views)93.6% → 54.7%87.3% → 86.5%90.4% → 89.3%95.6% → 92.5%
New Open Images photos (n = 2,505 + 2,505 views)83.8% → 44.5%75.5% → 73.5%84.5% → 78.9%90.7% → 84.4%

Woehrer 2026 was trained on clean photos, so this gap mostly reflects what each model was built for.

How it works

Diagram titled How RightWayUp works. A tilted city frame feeds a DINOv2 vision transformer fine-tuned for roll by ORTUS AI, which outputs a probability over all 360 degrees peaking at plus 37.7 degrees. Below, the decision path: high confidence gives an answer, unsure goes to the large model in Balanced and Pro, still unsure abstains.
From the launch video. Image: Poly Haven HDRI (CC0), view extracted and rotated by us; the ring and numbers are real RightWayUp output.

The backbone is a DINOv2 vision transformer (Apache-2.0, from Meta), fully fine-tuned for roll. It splits the frame into 14 × 14-pixel patches and reads each one as a token. The head outputs a probability for each of 360 one-degree bins. It is trained against a circular Gaussian target, so 359° and 1° count as neighbours. The peak is the angle, and the confidence is the probability within ±10° of the peak.

Training rotates images at random over the full 0 to 360° and crops them in a way that doesn’t depend on the angle, so the outline never gives the answer away. It also applies the simulated CCTV degradations described above. The released models were retrained so they don’t rely on the faint JPEG-grid and resampling traces that rotating a photo can leave, and every shipped file passes that check in its own runtime (details in the technical report).

Pico, Nano and Fast are single ViT-S/14 models. Nano reads images at 112 px, and Pico is Nano fine-tuned to read them at 70 px. Fast reads them at 224 px and was distilled from Max. Max is a ViT-L/14 that reads every image at 280 px. Balanced and Pro are cascades: Fast answers first, and its least-confident images go to Max (about 10% of calibration images for Balanced and 20% for Pro). If the answer is still uncertain, the tier abstains. Routing and abstention thresholds were fixed on calibration data and are not tuned per deployment. The routed share depends on the workload; hard images such as thermal frames are routed more often, which makes the cascades slower there.

Knowing when not to answer

Some images have no defined “up” (e.g. open sky, sand, a bare wall or a straight-down aerial view), so below a confidence threshold the model abstains rather than guess. We fixed the thresholds on separate calibration data, where the standard point answers about 90% of images, then applied them unchanged to every test. A strict point allows at most 1% wrong on calibration data, and abstention can be switched off. Each file format (ONNX FP32, FP16 and INT8, and Core ML) has its own thresholds, fitted on the same calibration data.

“I think abstention matters more than the headline number, because for a camera check a confident wrong answer costs more than a skipped frame.”

Vladimir Ilyash, CTO and co-founder, ORTUS AI

Card titled Knows when there is no up, with the line: calibrated abstention, thresholds fixed on calibration data only, per file format, then applied unchanged. Top row: a canal-side city view, a meadow below a mountain and a hotel room, each marked upright with the correction applied. Bottom row: blue sky, sand and a plain asphalt texture, each marked no defined up, abstains.
Poly Haven HDRIs (CC0), views extracted and rotated by us; real RightWayUp Max outputs.

Sealed final holdout, standard operating point. Each cell: share of images the tier answers / share wrong (more than 10° off) among the answered images.

Test set Max Pro Balanced Fast Nano
New COCO photos (n = 2,201)91% / 2.1%91% / 2.2%89% / 3.4%89% / 4.6%89% / 7.1%
Same photos, simulated CCTV degradation86% / 3.0%86% / 3.1%87% / 4.0%87% / 4.5%88% / 6.8%
New Open Images photos (n = 2,505)74% / 1.2%74% / 1.7%75% / 3.2%73% / 4.4%71% / 10.3%
Same photos, simulated CCTV degradation64% / 2.4%64% / 2.8%66% / 4.1%67% / 5.5%70% / 11.1%
Aalborg thermal camera (n = 150)80% / 12.5%84% / 19.8%83% / 27.2%74% / 31.5%77% / 50.0%

On unfamiliar photos the models answer less often rather than guess: on the new Open Images photos Max answers 74% of images and is wrong on 1.2% of those. Nano’s confidence is less dependable on photos (10.3% wrong among the Open Images photos it answers). Thermal cameras are the weak spot. On the unseen Aalborg thermal camera every tier in the table stays confident while wrong too often, and the strict point does not fix it. On thermal footage, use Max and don’t rely on the confidence alone.

Speed

We report two measurements separately. Single image (batch 1) is the median latency for one image processed on its own, which is what a per-camera check sees. Batched (batch N) is the sustained throughput when many images are processed together.

Pico, Nano and Fast take under 1 ms per image on a GPU: 0.55 ms, 0.55 ms and 0.68 ms for a single image (batch 1) with their one-image files on an NVIDIA RTX PRO 4500, TensorRT FP16. On an Apple M4 with Core ML FP16 on the Neural Engine they take 0.86 ms, 1.04 ms and 3.44 ms. On a laptop CPU (Intel Core i7-1260P, ONNX Runtime INT8, 4 threads) they take 3.21 ms, 7.73 ms and 29.7 ms, against 161 ms for Woehrer 2026 in FP32 on the same CPU. That makes Fast 5.4× and Pico 50× faster, INT8 against FP32.

Max is the accurate tier, not the fast one. For a single image it is slower than Woehrer 2026: 1.1 to 1.3× on the data-centre and workstation GPUs we measured and 1.6 to 5× on CPUs. Batched on a GPU, it gets through more images per second than Woehrer 2026, whose released graph takes one image at a time.

Bar chart titled Under 1 ms per image on a GPU. Real-time on a laptop CPU, too. RightWayUp Fast, single image (batch 1), median latency on a log scale: NVIDIA RTX PRO 4500, TensorRT FP16, one-image file, 0.68 ms; Apple M4 Mac mini, Core ML on the Neural Engine, 3.44 ms; Intel Core i7-1260P laptop, ONNX Runtime INT8, 4 threads, 29.7 ms; Woehrer 2026 on the same laptop CPU, ONNX Runtime FP32, 4 threads, 161 ms. A dashed line marks 33 ms, real time at 30 frames per second. Footer: 5.4 times faster than Woehrer 2026 on the same CPU, our INT8 against its FP32; Pico and Nano are faster still, 0.55 ms on the same GPU and 3.2 and 7.7 ms on the same laptop CPU; Max is the accurate tier, not the fast one.
RightWayUp Fast, single image (batch 1), median latency, log scale. GPU: TensorRT FP16, one-image file. Apple M4: Core ML on the Neural Engine. CPU: ONNX Runtime INT8, 4 threads. Woehrer 2026: ONNX Runtime FP32, 4 threads, on the same laptop CPU. The dashed line is 33 ms, one frame at 30 fps.

Single image (batch 1), median latency. GPUs run the batch-capable files; the one-image files of Pico, Nano and Fast are faster still (above). Balanced and Pro are derived from measured stage timings at the calibration routing share (10% / 20%); a workload that routes more images will be slower. Some rows are scaled to our earlier measurements on the same device; the method and every device we measured are in the benchmark results on GitHub.

Hardware and runtime Pico Nano Fast Balanced (derived) Pro (derived) Max Woehrer 2026
NVIDIA RTX PRO 4500 GPU, TensorRT FP160.75 ms0.95 ms1.25 ms1.77 ms2.29 ms5.19 ms4.04 msONNX Runtime CUDA FP32
NVIDIA RTX 5090 GPU, TensorRT FP160.68 ms0.86 ms0.98 ms1.48 ms1.98 ms4.99 ms3.73 msONNX Runtime CUDA FP32
NVIDIA L4 GPU, TensorRT FP160.88 ms1.62 ms1.26 ms2.18 ms3.09 ms9.15 ms7.28 msONNX Runtime CUDA FP32
Apple M4, Core ML FP16 on the Neural Engine0.86 ms1.04 ms3.44 ms8.78 ms14.1 ms53.5 ms138 msCPU FP32, 4 threads
Intel Core i7-1260P laptop CPU, ONNX Runtime INT8, 4 threads3.21 ms7.73 ms29.7 ms62.2 ms94.8 ms326 ms161 msFP32
Intel Xeon Gold 6530 server CPU, ONNX Runtime INT8, 4 threads7.43 ms11.5 ms24.3 ms44.2 ms64.1 ms199 ms123 msFP32
AMD EPYC 7663 server CPU, ONNX Runtime INT8, 4 threads12.9 ms18.6 ms42.4 ms95.8 ms149 ms534 ms150 msFP32

Batched (batch 64), sustained throughput in images per second, TensorRT FP16. Balanced and Pro are derived as above. Woehrer 2026’s released graph has a fixed batch of 1, so it has no batched figure; its single-image latency is in the table above.

GPU Pico Nano Fast Balanced (derived) Pro (derived) Max
NVIDIA RTX PRO 450055,950 img/s25,737 img/s6,483 img/s2,631 img/s1,650 img/s443 img/s
NVIDIA RTX 509062,071 img/s34,706 img/s10,013 img/s4,827 img/s3,180 img/s932 img/s
NVIDIA L431,845 img/s12,889 img/s2,381 img/s947 img/s591 img/s157 img/s

Batching amortises fixed per-call overhead and fills the GPU, so batched throughput is far higher than 1000 ÷ single-image latency. Every file format has its own abstention thresholds, fitted on the same calibration data. On CPU, INT8 is the recommended path.

Where it’s weaker

  • In straight-down (nadir) aerial images, “up” points out of the picture, so there is no “up” within the two-dimensional image. The model should abstain. Don’t use it to orient nadir drone footage.
  • Thermal cameras are the weakest case. On the unseen Aalborg thermal camera Max scores 83.3%, Nano and Fast do no better than Woehrer 2026, and the confidence of every tier we scored there is less reliable, so don’t rely on abstention for thermal footage. Most likely that is because thermal frames are a small part of the training data; fine-tuning on thermal footage should close much of the gap.
  • Text-heavy images and people lying down are among the hardest cases, with occasional 180° flips.
  • Nano and Fast trail Woehrer 2026 on clean COCO photos (87.3% and 90.4% against 93.6%), and Nano also on clean Open Images photos. They are built for scenes such as CCTV, rooms and streets, not close-up or top-down object photos. Pico gives up more accuracy for size and speed, and has development results only.
  • If your images are never far from level and you need full camera calibration, GeoCalib is the better choice.
  • It isn’t an IMU replacement, and we haven’t validated it on robots or vehicles.
  • Training photos are mostly Flickr consumer photography, and the CCTV training data comes from one US campus dataset.

Data provenance

Rotation labels look free, since you can rotate any upright image, but you still need to know that the image was upright to begin with and that you have the rights to use it. Max, Fast and Pico were trained on about 1.24 million distinct images (1,244,222 in 1,245,604 training rows), each with a recorded source and licence, from sources that permit commercial use: 1.22 million photos plus DIODE scans, MEVA CCTV frames and Poly Haven 3D renders. Nano’s final training came before the Open Images V7 and CommonCatalog expansion and used 705,075 training rows. For the 629,551 COCO, Open Images and CommonCatalog photos, we re-checked each photo’s current Flickr licence; PASS photos carry the licence record of the PASS dataset. No customer data was used.

  • PASS: Flickr photos without people; CC BY 4.0 dataset grant, with per-image CC BY 2.0 attribution.
  • COCO: CC BY photos and a small public-domain-like set, with the current Flickr licence re-checked per image.
  • Open Images (test subset and V7 train): CC BY photos and a small public-domain-like set, with the current Flickr licence re-checked per image, kept only where the recorded display orientation is upright and there is no EXIF rotation flag.
  • CommonCatalog: YFCC100M photos whose current Flickr licence is exactly CC BY 2.0.
  • DIODE: MIT.
  • MEVA: CC BY 4.0.
  • Poly Haven: CC0 assets, rendered by us. Our renders are published as the RightWayUp training renders dataset (CC BY 4.0).

We don’t redistribute the photos and frames, only our own renders. The published manifests trace every image to its source, with author credit where the licence requires it, so anyone can rebuild the set under the original terms. Near-duplicates of test and benchmark photos were excluded from training. The DINOv2 backbone is Apache-2.0. Meta doesn’t warrant the rights trail of its pretraining data, and our data card says so.

Use it

pip install rightwayup
rightwayup predict frame.jpg   # angle_cw, confidence, abstain
rightwayup fix photo.jpg       # writes an upright copy; leaves it alone if it abstains
from PIL import Image
from rightwayup import Orienter

o = Orienter(tier="max")           # downloads ONNX weights from the Hugging Face Hub
r = o.predict("frame.jpg")         # r.angle_cw, r.confidence, r.abstain, r.routed
img = Image.open("photo.jpg")
upright = o.correct(img, snap=90)  # rotated back to upright, snapped to the nearest 90°

angle_cw is the clockwise rotation of the image content, so rotating by that amount counter-clockwise restores it. Options for the tier (pico, nano, fast, balanced, pro or max), the abstention point (standard, strict or off), the device and batching are in the README on GitHub.

Licence and citation

Code and weights are released under Apache-2.0, including the Max tier, and commercial use is allowed. Keep the LICENSE and NOTICE files when redistributing. “ORTUS AI” and “RightWayUp” are trademarks and are not covered by the licence; please give modified or retrained models a different name.

If RightWayUp helps your product, research or project, we would appreciate a mention such as “Orientation by RightWayUp from ORTUS AI” with a link to this page, or a citation. This request adds no condition to the licence.

@misc{ortusai2026rightwayup,
  title  = {RightWayUp: Full-circle image roll estimation with calibrated abstention},
  author = {ORTUS AI},
  year   = {2026},
  url    = {https://github.com/ortusaitech/rightwayup}
}

Built for CHEQIT

We built RightWayUp for our CHEQIT camera-health work, to spot from a single frame that a camera has been knocked, rolled or remounted, or was mounted at the wrong angle in the first place, and to say by how many degrees. That saves maintenance visits and guesswork. CHEQIT is ORTUS AI’s on-prem monitoring for CCTV fleets.

Card titled Knocked camera? Know exactly how far. Left: a street camera feed tilted by a knock, with a measured roll of plus 26.6 degrees. Right: the corrected view with an illustrated alert reading camera rotated, CAM 07 turned 27 degrees clockwise from its baseline. Below, a roll timeline jumps above the plus or minus 10 degree alert threshold.
From the launch video. Static view: Poly Haven HDRI (CC0). The knock is simulated by us and the roll is measured per frame by RightWayUp Max. The CHEQIT alert card is an illustration.

If it fails on your images, we’d like to hear about it at hello@ortusai.io. If you need it adapted to thermal, fisheye or document imagery, or tuned for specific edge hardware, please reach out. We’re happy to work with you on it.

Credits

  • Roller-coaster footage in the video: “Huracan POV (Front Row)” and “Impulse POV” by Canobie Coaster, CC BY 3.0, via Wikimedia Commons (Huracan, Impulse). Modified: levelled by RightWayUp.
  • Action footage in the video: Pexels (Pexels License), Alex Moliski (mountain bike) and other Pexels creators (FPV and alpine coaster).
  • Scenes in the video and the stills on this page: Poly Haven HDRIs (CC0). Training-domain tile in the video: MEVA (Kitware / IARPA, CC BY 4.0).
  • Music and sound effects: generated with ElevenLabs on a paid plan that allows commercial use. ElevenLabs says its music model is trained on licensed stems and music, built in partnership with artists, labels and publishers. Font: Inter (SIL Open Font License).
  • Comparison models, each scored from its public release: Woehrer 2026 (code, paper, MIT licence); Deep-OAD, on the held-out sets only (code, scored via its public ONNX export). Also mentioned: GeoCalib (code).

All figures come from the RightWayUp 1.0 release weights. The sealed final holdout was opened once, after the weights, tiers and thresholds were frozen on 29 September 2026, and scored once; the development tests, which were also used to choose the models, were scored with the same weights. The held-out CCTV sets were used once before, on 24 September, to score our earlier pre-release models, and scored once more with the released models. Pico was chosen after the holdout was opened, so it has development results only. “Within 10°” is the share of images whose estimate is within 10° of the true angle.