andery /

Benchmarks

5 Sep 2026 · Pixel 9 Pro / Tensor G4 · Android 17 · ONNX Runtime 1.23.2

ModelsKeep Small + INT8Larger models showed no dependable quality gain.

SpeedTest CPU4 in the app~11× faster image inference; ~2× faster face embeddings.

PeopleReduce missed matches97/200 known queries rejected; 50 people split into 86 groups.

Measured encoder speedups on this phone, not whole-app speedups. Production models, runtime and thresholds are unchanged.

Image embeddings

Flickr8k · 176 photos · 880 caption queries
ModelTop-1 ↑Top-5 ↑Top-10 ↑Image ↓Text ↓Size ↓
UForm3 Smallcurrent77.16%95.11%98.30%28.8 ms4.3 ms60.6 MB
UForm3 Base76.48%95.45%98.41%96.0 ms7.2 ms125.3 MB
PE-S INT878.64%97.27%99.20%169.0 ms15.0 ms98.1 MB

Impact: PE-S adds 1.48 points of top-1 recall and 37.5 MB. Its 95% interval includes no gain (−3.30 to +6.25 points). Base is −0.68 points. Small remains the best fit in this test.

Top-k: paired photo appears in the first k results among 176 candidates. Times are full-pass medians on CPU4, excluding preprocessing. Base and PE-S reached light thermal throttling; timing ratios are not controlled.

Face embeddings & people

LFW · 600 pairs · 50 people / 300 grouping photos
ModelPair accuracy¹ ↑Known named² ↑Groups³Group F1⁴ ↑Size ↓
AuraFace INT8current82.7%103 / 20086 / 5090.0%65.7 MB
AuraFace FP3282.8%102 / 20087 / 5089.6%260.7 MB

Impact: FP32 adds 195 MB without improving named matching or grouping. INT8 rejected 97 known queries and split identities into extra groups. Detection and threshold calibration deserve attention before a bigger model.

¹ Same/different-person accuracy at 0.55. ² Correct matches at 0.65, with two clean enrollment photos per person. ³ Predicted groups / true identities. ⁴ B-cubed F1 combines group purity and completeness; five shuffled runs gave the same results. INT8 app policy and CPU4 produced identical embeddings.

No wrong identity matches in 200 known queries, no accepts in 200 unknown queries, and no incorrect group merges in this sample. This does not establish a zero false-match rate in larger galleries.

Pair thresholds, grouping and full-pass timings
ModelThresholdAccuracyTrue matchesFalse matches
AuraFace INT80.5582.7%196 / 3000 / 300
AuraFace INT80.6566.8%101 / 3000 / 300
AuraFace FP320.5582.8%197 / 3000 / 300
AuraFace FP320.6566.8%101 / 3000 / 300

All 600 pairs were evaluated. At the named threshold of 0.65, both models accepted 101 of 300 same-person pairs. No replacement threshold was established. Calibrate on separate validation data, then evaluate on a fresh held-out set.

INT8 grouping pair precision was 100%, pair recall 78.1%; FP32 recall was 77.5%. Named matching uses clean enrollment, not the UI naming workflow or enrollment from an impure cluster.

Full-pass median embedding times: INT8 CPU4 69.1 ms; INT8 app policy 226.1 ms; FP32 CPU4 393.5 ms. All 1,101 faces completed. These sequential passes reached light thermal throttling and do not provide a controlled speed ratio.

Face detection

WIDER FACE · 122 images · 1,866 valid faces
Decode limitFaces found ↑Recall ↑Precision ↑Median time ↓
512 px · current411 / 1,86622.0%81.2%38.4 ms
1024 px · experiment556 / 1,86629.8%87.0%239.9 ms

Impact: Larger input recovered 145 more faces, but still missed most in this difficult sample. At 512 px, recall was 81.6% for faces with both edges ≥32 px after downscaling. Small faces are the main gap.

ML Kit FAST 16.1.7, minimum face size 0.05; IoU ≥0.5 matching. This is sampled recall/precision, not official WIDER AP. The 1024 run had different power, focus and thermal conditions; its timing is not a controlled comparison.

Detection coverage and size breakdown

LFW: a face was detected in all 1,101 photos; 165 had multiple detections and the largest was selected. This measures coverage, not perfect localization. All selected faces had eye alignment. Median detection 35.6 ms; alignment 1.1 ms.

Minimum face edge at 512 pxFound / annotatedRecall
All valid faces411 / 186622.0%
Both edges ≥32 px93 / 11481.6%
Both edges ≥64 px36 / 3992.3%
Both edges ≥100 px16 / 1984.2%

Rows overlap. The 512 pass ran foreground on battery at thermal status 0; the 1024 pass ended charging, without window focus, at thermal status 1.

Runtime impact

Repeated short runs · median ms · lower is better
EncoderCPU4 · two sweepsXNNPACK · two sweepsApp policy · two sweeps
UForm3 Small image26.24 / 29.4354.63 / 54.50311.46 / 322.68
AuraFace INT854.32 / 53.08106.77 / 107.40107.32 / 106.92

Impact: CPU4 is the strongest lead: roughly 11× faster for the image encoder and 2× for face embeddings than NNAPI-first app policy. Validate sustained indexing, battery use and other phones before changing production.

Both sweeps at thermal status 0; five prepared inputs, eight warm-ups; 20 timed image calls / 30 face calls per case. Excludes decoding, alignment, model loading and database work. Accelerator selection does not imply full model offload.

Public datasets only; no personal gallery photos. Small samples on one physical phone. No full-gallery, battery, OCR or demographic evaluation; training overlap has not been audited.

Datasets, preprocessing and limits

Images: first 176 of 200 downloaded Flickr8k test images, five captions each. The last 24 were reserved for rejected static quantization. The evaluation data also informed exploratory choices, so it is not untouched final validation. Paired image-level bootstrap: 10,000 samples, seed 42. Base’s top-1 difference interval is −5.57 to +4.20 points; PE-S’s is −3.30 to +6.25.

Models: Small uses app-like bilinear square resizing to 224 px and 64 tokens; Base uses bicubic center crop at 224 px and its exported 77-token input. PE-S uses square bilinear 384 px inputs and dynamic per-channel INT8. A static QDQ export degraded retrieval and was rejected. PE native and desktop outputs differed (median cosine ~0.964 image / ~0.962 text); these results use native outputs.

Faces: first official LFW fold (300 positive / 300 negative pairs), not the complete ten-fold protocol. Grouping uses six photos each for 50 identities and five shuffled insertion orders. WIDER samples two images per validation event, seed 42. Actual Android ImageDecoder, ML Kit wrapper, two-eye alignment at 112 px, AuraFace embedder and PeopleService ran in a separate benchmark app.

Limits: No no-match image queries or thousands of distractors. LFW does not establish performance on children, aging, relatives, occlusion or demographic groups. Face grouping metrics are conditional on successful detection; this sample had no detection failures. Model sizes exclude runtime libraries. Full-pass timing conditions differ from the repeated short runtime sweeps.

Sources & pinned revisions