andery/field notesHow we measured
ON-DEVICE ML / EXPERIMENT 001

A bigger model.
A better gallery?

We tested three image-text models on a Pixel 9 Pro. For Andery's photo search, the smallest one still makes the most sense.

THE DECISION

Keep UForm3 Small.
Revisit the runtime.

PE-S gained 1.48 percentage points in top-1 retrieval, with an uncertainty interval that includes no gain. It also needed more storage and image compute. Base offered no measured improvement.

The stronger lead is CPU configuration. Four CPU threads ran Small's image encoder roughly 11× faster than NNAPI-first in controlled sweeps on this Pixel.

A runtime recommendation to validate in the app. Production inference is unchanged.

01 / RETRIEVAL

Accuracy has a cost.

All three models. The same candidate pool.
Every embedding computed on the phone.

Top-1 retrieval

Paired image ranked first. Higher is better.

UForm3 SmallCurrent
77.16%
UForm3 Base
76.48%
PE-S INT8
78.64%

Top-1 difference vs Small: Base −0.68 pp; PE-S +1.48 pp. Neither 95% interval excludes zero.

Native Android evaluation · CPU, four threads
ModelTop-1Top-5Top-10ImageTextModel size
UForm3 SmallCURRENT77.16%95.11%98.30%28.82 ms4.29 ms60.58 MB
UForm3 Base76.48%95.45%98.41%96.00 ms7.17 ms125.27 MB
PE-S INT878.64%97.27%99.20%168.95 ms15.00 ms98.06 MB

Image and text columns show median inference time. Model size includes both encoders, in decimal MB. Full-pass thermal status was 0 for Small and 1, light throttling, for Base and PE-S. These timings are descriptive, not a precisely controlled speed ratio.

02 / EXECUTION PROVIDERS

The useful speedup
was already on the phone.

In two controlled sweeps, CPU with four threads beat both XNNPACK and the app's NNAPI-first policy. NNAPI split the graph into many partitions. Selecting an accelerator did not put the whole model on it.

For Small, the image encoder took 26–29 ms on CPU4 versus 311–323 ms with NNAPI-first.

This measures encoder inference. It does not mean the entire gallery indexes eleven times faster. Other phones may need a different policy.

Small · image encoder

Median milliseconds per sweep · lower is better

CPU426.24 / 29.43 ms
Recommended to test
XNNPACK54.63 / 54.50 ms
8 provider threads
NNAPI-first311.46 / 322.68 ms
Current app policy

Two sweeps · 5 prepared inputs · 8 warm-ups
20 timed inferences per case · thermal status 0

03 / METHOD & LIMITS

What these numbers can tell us.

Download results JSON
01

A small, public photo pool

We downloaded the first 200 test rows from the Flickr8k mirror. The final comparison uses the first 176 images and their 880 captions. Each query ranks those 176 images by normalized cosine similarity.

The last 24 images were reserved for a rejected static quantization experiment. The evaluation set also informed exploratory choices, so it is not an untouched final validation set.

02

Native outputs, prepared inputs

Pixel 9 Pro, Tensor G4, Android 17. ONNX Runtime Android 1.23.2 with four CPU threads. Every final image and caption embedding came from the Android runner.

Inputs were prepared in advance. Timings exclude image decoding, resizing, tokenization, model loading, database work and face indexing. The phone was charging and unlocked.

03

Uncertainty matters

We resampled by image, keeping its five captions together, for 10,000 paired bootstrap samples with seed 42. Base's top-1 difference has a 95% interval of −5.57 to +4.20 percentage points. PE-S spans −3.30 to +6.25.

Both include zero. A larger gallery-specific evaluation is needed before claiming a dependable quality improvement.

04

Still outside the experiment

No personal photos were used or uploaded. This does not measure no-match queries, thousands of distractors, OCR, faces, battery energy, sustained full indexing or other Android chips.

PE's Android outputs differed numerically from desktop outputs. That cause remains unresolved. We report the full native evaluation rather than substituting higher desktop scores.

Model preparation and rejected experiments

Small. Existing UForm3 English Small ONNX assets, 64-token input. Quality preprocessing approximates the app's square bilinear resize at 224 pixels. Exact Android bitmap resampling was not tested.

Base. Official ONNX pair, bicubic resize and center crop at 224 pixels. The exported text input takes 77 tokens.

PE-S. PE-Core-S-16-384 from the timm port, official square bilinear preprocessing and mean/std 0.5. We exported FP32 ONNX, then applied dynamic per-channel INT8 to image MatMul/Gemm and text MatMul/Gemm/Gather. Operations outside that recipe retain their original types. The pair shrank from 356.64 MB to 98.06 MB.

A static QDQ image export degraded retrieval and was rejected. A first image run exhausted the Java heap because the runner preloaded every input; streaming inputs fixed it and the full pass succeeded. An activity-launch timeout was also corrected. Neither failure is counted as model performance.

A one-image, one-caption FP32 export check matched Torch to ONNX with cosine approximately 1.0. Across the final set, median Android-to-desktop PE cosine was about 0.964 for images and 0.962 for text. This limits confidence in treating this converted model as representative of every PE deployment.

04 / PROVENANCE

Follow the inputs.