ModelsKeep Small + INT8Larger models showed no dependable quality gain.
SpeedTest CPU4 in the app~11× faster image inference; ~2× faster face embeddings.
PeopleReduce missed matches97/200 known queries rejected; 50 people split into 86 groups.
Measured encoder speedups on this phone, not whole-app speedups. Production models, runtime and thresholds are unchanged.
Image embeddings
Flickr8k · 176 photos · 880 caption queries| Model | Top-1 ↑ | Top-5 ↑ | Top-10 ↑ | Image ↓ | Text ↓ | Size ↓ |
|---|---|---|---|---|---|---|
| UForm3 Smallcurrent | 77.16% | 95.11% | 98.30% | 28.8 ms | 4.3 ms | 60.6 MB |
| UForm3 Base | 76.48% | 95.45% | 98.41% | 96.0 ms | 7.2 ms | 125.3 MB |
| PE-S INT8 | 78.64% | 97.27% | 99.20% | 169.0 ms | 15.0 ms | 98.1 MB |
Impact: PE-S adds 1.48 points of top-1 recall and 37.5 MB. Its 95% interval includes no gain (−3.30 to +6.25 points). Base is −0.68 points. Small remains the best fit in this test.
Top-k: paired photo appears in the first k results among 176 candidates. Times are full-pass medians on CPU4, excluding preprocessing. Base and PE-S reached light thermal throttling; timing ratios are not controlled.
Face embeddings & people
LFW · 600 pairs · 50 people / 300 grouping photos| Model | Pair accuracy¹ ↑ | Known named² ↑ | Groups³ | Group F1⁴ ↑ | Size ↓ |
|---|---|---|---|---|---|
| AuraFace INT8current | 82.7% | 103 / 200 | 86 / 50 | 90.0% | 65.7 MB |
| AuraFace FP32 | 82.8% | 102 / 200 | 87 / 50 | 89.6% | 260.7 MB |
Impact: FP32 adds 195 MB without improving named matching or grouping. INT8 rejected 97 known queries and split identities into extra groups. Detection and threshold calibration deserve attention before a bigger model.
¹ Same/different-person accuracy at 0.55. ² Correct matches at 0.65, with two clean enrollment photos per person. ³ Predicted groups / true identities. ⁴ B-cubed F1 combines group purity and completeness; five shuffled runs gave the same results. INT8 app policy and CPU4 produced identical embeddings.
No wrong identity matches in 200 known queries, no accepts in 200 unknown queries, and no incorrect group merges in this sample. This does not establish a zero false-match rate in larger galleries.
Pair thresholds, grouping and full-pass timings
| Model | Threshold | Accuracy | True matches | False matches |
|---|---|---|---|---|
| AuraFace INT8 | 0.55 | 82.7% | 196 / 300 | 0 / 300 |
| AuraFace INT8 | 0.65 | 66.8% | 101 / 300 | 0 / 300 |
| AuraFace FP32 | 0.55 | 82.8% | 197 / 300 | 0 / 300 |
| AuraFace FP32 | 0.65 | 66.8% | 101 / 300 | 0 / 300 |
All 600 pairs were evaluated. At the named threshold of 0.65, both models accepted 101 of 300 same-person pairs. No replacement threshold was established. Calibrate on separate validation data, then evaluate on a fresh held-out set.
INT8 grouping pair precision was 100%, pair recall 78.1%; FP32 recall was 77.5%. Named matching uses clean enrollment, not the UI naming workflow or enrollment from an impure cluster.
Full-pass median embedding times: INT8 CPU4 69.1 ms; INT8 app policy 226.1 ms; FP32 CPU4 393.5 ms. All 1,101 faces completed. These sequential passes reached light thermal throttling and do not provide a controlled speed ratio.
Face detection
WIDER FACE · 122 images · 1,866 valid faces| Decode limit | Faces found ↑ | Recall ↑ | Precision ↑ | Median time ↓ |
|---|---|---|---|---|
| 512 px · current | 411 / 1,866 | 22.0% | 81.2% | 38.4 ms |
| 1024 px · experiment | 556 / 1,866 | 29.8% | 87.0% | 239.9 ms |
Impact: Larger input recovered 145 more faces, but still missed most in this difficult sample. At 512 px, recall was 81.6% for faces with both edges ≥32 px after downscaling. Small faces are the main gap.
ML Kit FAST 16.1.7, minimum face size 0.05; IoU ≥0.5 matching. This is sampled recall/precision, not official WIDER AP. The 1024 run had different power, focus and thermal conditions; its timing is not a controlled comparison.
Detection coverage and size breakdown
LFW: a face was detected in all 1,101 photos; 165 had multiple detections and the largest was selected. This measures coverage, not perfect localization. All selected faces had eye alignment. Median detection 35.6 ms; alignment 1.1 ms.
| Minimum face edge at 512 px | Found / annotated | Recall |
|---|---|---|
| All valid faces | 411 / 1866 | 22.0% |
| Both edges ≥32 px | 93 / 114 | 81.6% |
| Both edges ≥64 px | 36 / 39 | 92.3% |
| Both edges ≥100 px | 16 / 19 | 84.2% |
Rows overlap. The 512 pass ran foreground on battery at thermal status 0; the 1024 pass ended charging, without window focus, at thermal status 1.
Runtime impact
Repeated short runs · median ms · lower is better| Encoder | CPU4 · two sweeps | XNNPACK · two sweeps | App policy · two sweeps |
|---|---|---|---|
| UForm3 Small image | 26.24 / 29.43 | 54.63 / 54.50 | 311.46 / 322.68 |
| AuraFace INT8 | 54.32 / 53.08 | 106.77 / 107.40 | 107.32 / 106.92 |
Impact: CPU4 is the strongest lead: roughly 11× faster for the image encoder and 2× for face embeddings than NNAPI-first app policy. Validate sustained indexing, battery use and other phones before changing production.
Both sweeps at thermal status 0; five prepared inputs, eight warm-ups; 20 timed image calls / 30 face calls per case. Excludes decoding, alignment, model loading and database work. Accelerator selection does not imply full model offload.
Data & method
Public datasets only; no personal gallery photos. Small samples on one physical phone. No full-gallery, battery, OCR or demographic evaluation; training overlap has not been audited.
Datasets, preprocessing and limits
Images: first 176 of 200 downloaded Flickr8k test images, five captions each. The last 24 were reserved for rejected static quantization. The evaluation data also informed exploratory choices, so it is not untouched final validation. Paired image-level bootstrap: 10,000 samples, seed 42. Base’s top-1 difference interval is −5.57 to +4.20 points; PE-S’s is −3.30 to +6.25.
Models: Small uses app-like bilinear square resizing to 224 px and 64 tokens; Base uses bicubic center crop at 224 px and its exported 77-token input. PE-S uses square bilinear 384 px inputs and dynamic per-channel INT8. A static QDQ export degraded retrieval and was rejected. PE native and desktop outputs differed (median cosine ~0.964 image / ~0.962 text); these results use native outputs.
Faces: first official LFW fold (300 positive / 300 negative pairs), not the complete ten-fold protocol. Grouping uses six photos each for 50 identities and five shuffled insertion orders. WIDER samples two images per validation event, seed 42. Actual Android ImageDecoder, ML Kit wrapper, two-eye alignment at 112 px, AuraFace embedder and PeopleService ran in a separate benchmark app.
Limits: No no-match image queries or thousands of distractors. LFW does not establish performance on children, aging, relatives, occlusion or demographic groups. Face grouping metrics are conditional on successful detection; this sample had no detection failures. Model sizes exclude runtime libraries. Full-pass timing conditions differ from the repeated short runtime sweeps.