Downloads · 30 days
0
resoajoe/upscale-factor-nano
upscale-factor-nano is a image classification model from resoajoe. Use it when you need a label for an image. The card lists the license as mit.
46,899 parameters. 187 KB. Was this image upscaled, and by how much?
Downloads · 30 days
0
Access
Public
Updated Sep 3, 2026
Repo size
391 KB
Likes
0
Public
Click a slice to open those files.
.pt199 KB · 49%
From the Hugging Face model README
46,899 parameters. 187 KB. Was this image upscaled, and by how much?
Given a 64×64 patch, predicts whether it was resampled up from a smaller original by 2×, 3×, or 4× — the signature of a "4K" stream that is really 720p wearing a bigger container.
Domain measured / deployment domain tested: measured on COCO val2017 photographs with real bicubic/bilinear resampling; deployment domain: no GAN or learned upscaler, and no camera ISP upscaling, tested. (Fifth line of the card standard, added 2026-09-02: a number is only as good as the domain it was measured in.)
Trained on real photographs (COCO val2017) with real resampling. Split by source image.
| task | nano (held out) | best scalar (in-sample) | chance |
|---|---|---|---|
| upscale factor — 2× / 3× / 4× | 0.980 | 0.489 | 0.344 |
| upscaled at all, yes/no | 0.977 | 0.847 | 0.514 |
The factor task is the one worth having. Energy loss alone cannot answer "by how much" — every tested scalar (mean, std, Laplacian variance, HF ratio, gradient, entropy) lands between 0.357 and 0.489, because the answer lives in the periodicity interpolation leaves behind, not in how much detail is missing. The model beats the best optimistically-fitted scalar by +0.491.
A prediction made before running this was wrong and is recorded here for honesty: the binary task was expected to fall to a scalar, since upscaling removes high frequencies much as blur does. It did not — 0.977 vs 0.847. Upscaling leaves structure a single statistic cannot capture even for the coarse question.
Everything above is COCO — one corpus, one kind of imagery. Tested on 1,002 patches from a Logitech BRIO watching a real office: a different sensor, a different image pipeline with its own denoising and sharpening, and frames that passed through the camera's own MJPG compression on the way out.
| accuracy | 2x | 3x | 4x | |
|---|---|---|---|---|
| COCO photographs (held out) | 0.982 | 0.981 | 0.964 | 1.000 |
| Logitech BRIO, real office | 0.876 | 0.757 | 0.877 | 0.997 |
It transfers, losing about 10 points. The loss is entirely concentrated in the 2x class (0.981 → 0.757) while 4x is untouched (1.000 → 0.997), which is the physically expected shape: 2x upscaling leaves the subtlest periodic signature and lossy compression attenuates exactly those high frequencies first. This card previously listed "heavy compression after upscaling" as an untested risk; it is now measured, and the number to carry is 2x detection degrades first.
Flat patches (std < 6) were skipped, since sky and blank wall contain nothing to resample.
Limitation of this test: the BRIO offers a maximum of 1920x1080 over V4L2, so whether those frames are sensor-native or internally rescaled from a larger array could not be determined here. If they are rescaled, the real-world number is if anything pessimistic.
For: deciding whether to spend decode budget on a stream that has no real detail, media triage, dataset hygiene (finding upscaled images in a training corpus), and quality auditing on cheap hardware.
Not for:
import cv2, numpy as np, onnxruntime as ort
FACTORS = [2, 3, 4]
sess = ort.InferenceSession("upscale_factor.onnx", providers=["CPUExecutionProvider"])
img = cv2.imread("frame.png", cv2.IMREAD_GRAYSCALE)
p = img[y:y+64, x:x+64].astype(np.float32) # a NATIVE 64x64 crop -- do not resize first
p = (p - p.mean()) / (p.std() + 1e-8)
print(FACTORS[int(sess.run(None, {"input": p[None,None]})[0][0].argmax())])
Take a native crop; never resize the image to 64×64 first. Resizing is itself a resampling operation and destroys the very signature the model reads. This is the single easiest way to get meaningless output.
Patches are cheap — run several across the frame and take the majority vote.
ONNX vs PyTorch, both CPU, 256 inputs: max relative logit difference 2.8e-07, 100% argmax agreement.
Every margin quoted here is against a stated baseline, because a margin without one is not a measurement. The baseline is the best single-threshold classifier over ten cheap statistics, fitted optimistically:
mean · std · lapvar · hf (high-frequency energy ratio) · grad (Sobel magnitude) ·
entropy · centre_edge · radial_slope · row_fft_peak · col_fft_peak
The last four are spatially aware, added after an earlier six-statistic baseline — all global aggregates — was found to systematically overstate model value on spatially structured tasks. A baseline that cannot see where anything is loses to a CNN by default. On one test task that flaw inflated an apparent margin from +0.060 to +0.261.
Two questions are asked with it, and they disagree:
Where this card quotes a single scalar figure without qualification, it is the in-sample one.
COCO val2017, a public dataset. No personal data involved.