Research: licensable face detection and recognition models #31

Open
opened 2026-09-24 14:45:09 +00:00 by kayg · 5 comments
Owner

Owner decision (round 13, P5): faces ship after CLIP/map, as soon as a model with an acceptable license exists. Immich's InsightFace weights are not redistributable for third-party or commercial use.

Research deliverable (web search, verify licenses at the source, as of the research date): candidate face detection + recognition (embedding) models in ONNX with permissive licenses (Apache-2.0/MIT/BSD or explicit commercial permission), their accuracy on standard benchmarks (e.g. LFW/IJB), speed on CPU (arm64 + amd64), model size, and provenance of training data. Recommend one pair, list risks. Output as a comment on this issue; no code.

Context for the owning job

  • Repo: kayg/calternal (~/Developer/calternal). Read CLAUDE.md, CONTEXT.md and docs/DESIGN.md first; this issue's section is cited below.
  • Owner rules that always apply: file over app (plain files are the truth, the DB is an index); the server is the single writer; data loss is unacceptable; performance first, never at the cost of finesse; UI is the calternal.js design system (copy components verbatim, compare side by side with calternal.js reference screenshots; Claude does visual review); never ship sample/mock data; atomic commits; adversarial testing after API work; good enough, not perfect (merge blockers: crash/DoS, data loss, security, sync collisions).
  • Comment on this issue when you start (branch, base SHA), on each finding, when blocked, and when finished (head SHA + gate output). Never close it.
Owner decision (round 13, P5): faces ship after CLIP/map, as soon as a model with an acceptable license exists. Immich's InsightFace weights are **not** redistributable for third-party or commercial use. Research deliverable (web search, verify licenses at the source, as of the research date): candidate face detection + recognition (embedding) models in ONNX with permissive licenses (Apache-2.0/MIT/BSD or explicit commercial permission), their accuracy on standard benchmarks (e.g. LFW/IJB), speed on CPU (arm64 + amd64), model size, and provenance of training data. Recommend one pair, list risks. Output as a comment on this issue; no code. ## Context for the owning job - Repo: kayg/calternal (~/Developer/calternal). Read CLAUDE.md, CONTEXT.md and docs/DESIGN.md first; this issue's section is cited below. - Owner rules that always apply: file over app (plain files are the truth, the DB is an index); the server is the single writer; data loss is unacceptable; performance first, never at the cost of finesse; UI is the calternal.js design system (copy components verbatim, compare side by side with calternal.js reference screenshots; Claude does visual review); never ship sample/mock data; atomic commits; adversarial testing after API work; good enough, not perfect (merge blockers: crash/DoS, data loss, security, sync collisions). - Comment on this issue when you start (branch, base SHA), on each finding, when blocked, and when finished (head SHA + gate output). Never close it.
Author
Owner

Research started in job/research-faces at base 57118d9648582e682f0a0e1997fc8ad9f84bab35. I am checking source licenses, benchmark evidence, model size, CPU performance evidence, and training-data provenance for ONNX face detection and recognition candidates. This is a research-only job; the worktree remains unchanged.

Research started in `job/research-faces` at base `57118d9648582e682f0a0e1997fc8ad9f84bab35`. I am checking source licenses, benchmark evidence, model size, CPU performance evidence, and training-data provenance for ONNX face detection and recognition candidates. This is a research-only job; the worktree remains unchanged.
Author
Owner

Initial source check: YuNet (face_detection_yunet_2023mar.onnx) and SFace (face_recognition_sface_2021dec.onnx) are the strongest small ONNX pair I found. OpenCV Zoo says every YuNet model-directory file is MIT-licensed and every SFace model-directory file is Apache-2.0; the OpenCV face tutorial lists 338 KB and 36.9 MB respectively. The tutorial reports YuNet WIDER Face validation AP of 0.830 / 0.824 / 0.708 (easy / medium / hard), and SFace verification accuracy of 99.60% LFW, 93.95% CALFW, 91.05% CPLFW, 94.90% AgeDB-30 and 94.80% CFP-FP.

Sources: YuNet model card, SFace model card, OpenCV model sizes and benchmark table.

Important provenance caveat: the SFace card maps the weight to the SFace loss and MobileFaceNet, but not to a specific dataset. The original SFace repository mentions experiments on CASIA-WebFace, VGGFace2 and MS-Celeb-1M without mapping the exact OpenCV ONNX weight. An OpenCV Zoo issue asking about that exact file/checksum and commercial permission has been open since 2026-07-22 without an upstream answer: issue #313. I will treat SFace as a candidate pending provenance confirmation, not as fully cleared for redistribution.

Initial source check: YuNet (`face_detection_yunet_2023mar.onnx`) and SFace (`face_recognition_sface_2021dec.onnx`) are the strongest small ONNX pair I found. OpenCV Zoo says every YuNet model-directory file is MIT-licensed and every SFace model-directory file is Apache-2.0; the OpenCV face tutorial lists 338 KB and 36.9 MB respectively. The tutorial reports YuNet WIDER Face validation AP of 0.830 / 0.824 / 0.708 (easy / medium / hard), and SFace verification accuracy of 99.60% LFW, 93.95% CALFW, 91.05% CPLFW, 94.90% AgeDB-30 and 94.80% CFP-FP. Sources: [YuNet model card](https://github.com/opencv/opencv_zoo/blob/main/models/face_detection_yunet/README.md), [SFace model card](https://github.com/opencv/opencv_zoo/blob/main/models/face_recognition_sface/README.md), [OpenCV model sizes and benchmark table](https://github.com/opencv/opencv/blob/5.x/doc/tutorials/dnn/dnn_face/dnn_face.markdown). Important provenance caveat: the SFace card maps the weight to the SFace loss and MobileFaceNet, but not to a specific dataset. The original SFace repository mentions experiments on CASIA-WebFace, VGGFace2 and MS-Celeb-1M without mapping the exact OpenCV ONNX weight. An OpenCV Zoo issue asking about that exact file/checksum and commercial permission has been open since 2026-07-22 without an upstream answer: [issue #313](https://github.com/opencv/opencv_zoo/issues/313). I will treat SFace as a candidate pending provenance confirmation, not as fully cleared for redistribution.
Author
Owner

CPU performance evidence (published CPU-only OpenCV DNN timings; no GPU):

  • OpenCV Zoo's benchmark README, OpenCV 4.10.0, batch 1: Intel Core i7-12700K x86-64 averaged 0.69 ms per YuNet call at 160×120 and 5.09 ms per SFace call at 150×150. Raspberry Pi 4B (Cortex-A72 / ARMv8 64-bit) averaged 6.23 ms and 68.82 ms at the same respective input sizes. These are small benchmark inputs, not full photo-size detector costs. Benchmark protocol/results
  • A separate 2026 third-party OpenCV 4.10.0 measurement used Raspberry Pi 5 Cortex-A76 (arm64) and Intel Core i7-13700K (x86-64), with 100-run medians. At 640×480, YuNet detection was 50.79 ms (Pi 5) / 3.29 ms (Intel); SFace feature extraction was 23.44 ms / 4.16 ms. At 1280×720, detection was 179.25 ms / 10.96 ms. The article publishes its measurement code and fixes power/frequency settings; it is from a commercial SDK vendor, so treat it as an external reproduction, not an independent lab benchmark. Conditions and results

These results indicate feasible background CPU indexing at reduced resolution, with detection as the cost driver. They do not establish calternal latency: Rust/ONNX Runtime, target CPU, image dimensions, thread settings and detector preprocessing all need measurement in the actual server path. Cache embeddings as derived data; do not recalculate them on every photo view.

CPU performance evidence (published CPU-only OpenCV DNN timings; no GPU): * OpenCV Zoo's benchmark README, OpenCV 4.10.0, batch 1: Intel Core i7-12700K x86-64 averaged 0.69 ms per YuNet call at 160×120 and 5.09 ms per SFace call at 150×150. Raspberry Pi 4B (Cortex-A72 / ARMv8 64-bit) averaged 6.23 ms and 68.82 ms at the same respective input sizes. These are small benchmark inputs, not full photo-size detector costs. [Benchmark protocol/results](https://github.com/opencv/opencv_zoo/blob/main/benchmark/README.md) * A separate 2026 third-party OpenCV 4.10.0 measurement used Raspberry Pi 5 Cortex-A76 (arm64) and Intel Core i7-13700K (x86-64), with 100-run medians. At 640×480, YuNet detection was 50.79 ms (Pi 5) / 3.29 ms (Intel); SFace feature extraction was 23.44 ms / 4.16 ms. At 1280×720, detection was 179.25 ms / 10.96 ms. The article publishes its measurement code and fixes power/frequency settings; it is from a commercial SDK vendor, so treat it as an external reproduction, not an independent lab benchmark. [Conditions and results](https://tech.swallow-incubate.com/liveness/cortex-a76-opencv-benchmark/) These results indicate feasible background CPU indexing at reduced resolution, with detection as the cost driver. They do not establish calternal latency: Rust/ONNX Runtime, target CPU, image dimensions, thread settings and detector preprocessing all need measurement in the actual server path. Cache embeddings as derived data; do not recalculate them on every photo view.
Author
Owner

Research result — 2026-09-24

Recommendation

Use YuNet detection + SFace embeddings as the first candidate pair. Both are ONNX models with explicit permissive model-directory license statements, and CPU measurements exist for x86-64 and arm64. Treat SFace as provisional pending model-data provenance review: Apache-2.0 is stated for all files in its OpenCV Zoo directory, but the exact training set behind the released ONNX file is not identified. If the owner requires confirmed training-data rights before accepting a model, I did not find a pair whose public sources settle both model and dataset provenance in this research pass.

Models

Role / exact artifact License stated for model files Size Accuracy and provenance
Detector: face_detection_yunet_2023mar.onnx (OpenCV Zoo YuNet) MIT; the model directory README says all files are MIT. 338 KB in the OpenCV face DNN tutorial. WIDER Face validation AP is reported as 0.830 / 0.824 / 0.708 (easy / medium / hard) in the OpenCV tutorial; YuNet's model page says 0.834 / 0.824 / 0.708. The same page's tools/eval table reports 0.8844 / 0.8656 / 0.7503. These published numbers disagree, so they should not be combined; reproduce one fixed evaluation protocol before treating the higher table as the model's expected accuracy. The model card links the YuNet training project, whose documentation names WIDER Face as the training data. The exact 2023mar training manifest and dataset-rights trail are not pinned in the model card.
Recognizer: face_recognition_sface_2021dec.onnx (OpenCV Zoo SFace) Apache-2.0; the model directory README says all files are Apache-2.0. 36.9 MB in the OpenCV face DNN tutorial. OpenCV reports verification accuracy: LFW 99.60%, CALFW 93.95%, CPLFW 91.05%, AgeDB-30 94.90%, CFP-FP 94.80%. The model card identifies MobileFaceNet trained with the SFace loss. The original SFace repo reports work on CASIA-WebFace, VGGFace2 and MS-Celeb-1M but does not map the exact OpenCV ONNX weights to a dataset or mixture.

Sources: YuNet model card, YuNet training project, SFace model card/license, OpenCV face model sizes and standard benchmark results, original SFace project.

CPU performance

OpenCV Zoo's published OpenCV 4.10.0 CPU-only DNN benchmark (batch 1) measured these per-call means:

Hardware YuNet, 160×120 SFace, 150×150
Intel Core i7-12700K, x86-64 0.69 ms 5.09 ms
Raspberry Pi 4B, Cortex-A72 / arm64 6.23 ms 68.82 ms

The detector input is only 160×120, so these are not full-resolution photo costs. A separate 2026 third-party OpenCV 4.10.0 test, with 100-run medians and the same models, reports more useful 640×480 figures: YuNet 50.79 ms on Raspberry Pi 5 Cortex-A76 (arm64) and 3.29 ms on Intel i7-13700K (x86-64); SFace feature extraction 23.44 ms and 4.16 ms respectively. At 1280×720, detection was 179.25 ms and 10.96 ms. The third-party article publishes its test code and conditions, but is from a commercial SDK vendor. These numbers are directional; calternal's Rust/ONNX Runtime backend, CPUs, thread settings and image pipeline must be benchmarked directly. OpenCV Zoo benchmark protocol/results; 640×480 / 1280×720 comparison and conditions.

Risks and handling

  1. SFace provenance is the blocking compliance uncertainty. Its OpenCV Zoo directory explicitly says Apache-2.0 for all files, but neither the model card nor conversion record ties this exact weight to the original repo's named datasets. A public request asking about the exact file/checksum and commercial use has remained open since 2026-07-22 without an upstream answer: OpenCV Zoo issue #313. Ask the contributor/maintainer to confirm the training set and whether its terms permit these released weights to be redistributed and used commercially. Keep the model artifact hash and license notice pinned if accepted.
  2. Do not substitute InsightFace packs as an easy alternative. InsightFace says its library code is MIT, but its pretrained model packs are for non-commercial research unless separately licensed: official model-zoo license note.
  3. Verification accuracy is not photo-library clustering accuracy. LFW and related figures measure pair verification under fixed protocols; they do not predict false merges across a user's 100k+ photo library. Calibrate thresholds on consented, representative photo data; make merges/rejections reversible and user-controlled. The published docs do not report subgroup results for this exact export.
  4. Resolution drives detection cost and misses. The separate CPU measurement shows 640×480 → 1280×720 adds about 3.5× YuNet time on A76. YuNet's card says its training scheme targets faces about 10–300 pixels. Use a measured downscaled scan plus original-resolution face crops, and test small/occluded faces.
  5. Identity vectors are sensitive derived data. Keep face vectors in the rebuildable Index/cache, not the Home; persist only human decisions (cluster name, merge, rejection) in .calternal/ collection files, as DESIGN.md §4 requires. Make processing opt-in and local. Face matching can fall within GDPR biometric-data rules when it enables unique identification; assess the actual product use and applicable legal basis before release (GDPR Article 4(14) and Article 9).
  6. Pin the artifact/runtime contract. OpenCV Zoo says face_detection_yunet_2023mar.onnx has fixed input dimensions and its 2026may variant is a dynamic-shape re-export. Pin a specific ONNX file, hash, license and preprocessing contract; verify compatibility with the selected ONNX Runtime version before shipping.

The design fit is good: both vectors are derived data; only the person's name and user merge/reject choices need durable collection-file storage. My recommendation remains YuNet + SFace, with SFace's training-data confirmation as an explicit pre-release gate.

## Research result — 2026-09-24 ### Recommendation Use **YuNet detection + SFace embeddings** as the first candidate pair. Both are ONNX models with explicit permissive model-directory license statements, and CPU measurements exist for x86-64 and arm64. Treat SFace as **provisional pending model-data provenance review**: Apache-2.0 is stated for all files in its OpenCV Zoo directory, but the exact training set behind the released ONNX file is not identified. If the owner requires confirmed training-data rights before accepting a model, I did not find a pair whose public sources settle both model and dataset provenance in this research pass. ### Models | Role / exact artifact | License stated for model files | Size | Accuracy and provenance | |---|---|---:|---| | Detector: `face_detection_yunet_2023mar.onnx` (OpenCV Zoo YuNet) | MIT; the model directory README says all files are MIT. | 338 KB in the OpenCV face DNN tutorial. | WIDER Face validation AP is reported as 0.830 / 0.824 / 0.708 (easy / medium / hard) in the OpenCV tutorial; YuNet's model page says 0.834 / 0.824 / 0.708. The same page's tools/eval table reports 0.8844 / 0.8656 / 0.7503. These published numbers disagree, so they should not be combined; reproduce one fixed evaluation protocol before treating the higher table as the model's expected accuracy. The model card links the YuNet training project, whose documentation names WIDER Face as the training data. The exact `2023mar` training manifest and dataset-rights trail are not pinned in the model card. | | Recognizer: `face_recognition_sface_2021dec.onnx` (OpenCV Zoo SFace) | Apache-2.0; the model directory README says all files are Apache-2.0. | 36.9 MB in the OpenCV face DNN tutorial. | OpenCV reports verification accuracy: LFW 99.60%, CALFW 93.95%, CPLFW 91.05%, AgeDB-30 94.90%, CFP-FP 94.80%. The model card identifies MobileFaceNet trained with the SFace loss. The original SFace repo reports work on CASIA-WebFace, VGGFace2 and MS-Celeb-1M but does not map the exact OpenCV ONNX weights to a dataset or mixture. | Sources: [YuNet model card](https://github.com/opencv/opencv_zoo/blob/main/models/face_detection_yunet/README.md), [YuNet training project](https://github.com/ShiqiYu/libfacedetection.train), [SFace model card/license](https://github.com/opencv/opencv_zoo/blob/main/models/face_recognition_sface/README.md), [OpenCV face model sizes and standard benchmark results](https://github.com/opencv/opencv/blob/5.x/doc/tutorials/dnn/dnn_face/dnn_face.markdown), [original SFace project](https://github.com/zhongyy/SFace). ### CPU performance OpenCV Zoo's published OpenCV 4.10.0 CPU-only DNN benchmark (batch 1) measured these per-call means: | Hardware | YuNet, 160×120 | SFace, 150×150 | |---|---:|---:| | Intel Core i7-12700K, x86-64 | 0.69 ms | 5.09 ms | | Raspberry Pi 4B, Cortex-A72 / arm64 | 6.23 ms | 68.82 ms | The detector input is only 160×120, so these are not full-resolution photo costs. A separate 2026 third-party OpenCV 4.10.0 test, with 100-run medians and the same models, reports more useful 640×480 figures: YuNet 50.79 ms on Raspberry Pi 5 Cortex-A76 (arm64) and 3.29 ms on Intel i7-13700K (x86-64); SFace feature extraction 23.44 ms and 4.16 ms respectively. At 1280×720, detection was 179.25 ms and 10.96 ms. The third-party article publishes its test code and conditions, but is from a commercial SDK vendor. These numbers are directional; calternal's Rust/ONNX Runtime backend, CPUs, thread settings and image pipeline must be benchmarked directly. [OpenCV Zoo benchmark protocol/results](https://github.com/opencv/opencv_zoo/blob/main/benchmark/README.md); [640×480 / 1280×720 comparison and conditions](https://tech.swallow-incubate.com/liveness/cortex-a76-opencv-benchmark/). ### Risks and handling 1. **SFace provenance is the blocking compliance uncertainty.** Its OpenCV Zoo directory explicitly says Apache-2.0 for all files, but neither the model card nor conversion record ties this exact weight to the original repo's named datasets. A public request asking about the exact file/checksum and commercial use has remained open since 2026-07-22 without an upstream answer: [OpenCV Zoo issue #313](https://github.com/opencv/opencv_zoo/issues/313). Ask the contributor/maintainer to confirm the training set and whether its terms permit these released weights to be redistributed and used commercially. Keep the model artifact hash and license notice pinned if accepted. 2. **Do not substitute InsightFace packs as an easy alternative.** InsightFace says its library code is MIT, but its pretrained model packs are for non-commercial research unless separately licensed: [official model-zoo license note](https://github.com/deepinsight/insightface/blob/master/python-package/docs/model_zoo.md). 3. **Verification accuracy is not photo-library clustering accuracy.** LFW and related figures measure pair verification under fixed protocols; they do not predict false merges across a user's 100k+ photo library. Calibrate thresholds on consented, representative photo data; make merges/rejections reversible and user-controlled. The published docs do not report subgroup results for this exact export. 4. **Resolution drives detection cost and misses.** The separate CPU measurement shows 640×480 → 1280×720 adds about 3.5× YuNet time on A76. YuNet's card says its training scheme targets faces about 10–300 pixels. Use a measured downscaled scan plus original-resolution face crops, and test small/occluded faces. 5. **Identity vectors are sensitive derived data.** Keep face vectors in the rebuildable Index/cache, not the Home; persist only human decisions (cluster name, merge, rejection) in `.calternal/` collection files, as DESIGN.md §4 requires. Make processing opt-in and local. Face matching can fall within GDPR biometric-data rules when it enables unique identification; assess the actual product use and applicable legal basis before release ([GDPR Article 4(14) and Article 9](https://eur-lex.europa.eu/eli/reg/2016/679)). 6. **Pin the artifact/runtime contract.** OpenCV Zoo says `face_detection_yunet_2023mar.onnx` has fixed input dimensions and its `2026may` variant is a dynamic-shape re-export. Pin a specific ONNX file, hash, license and preprocessing contract; verify compatibility with the selected ONNX Runtime version before shipping. The design fit is good: both vectors are derived data; only the person's name and user merge/reject choices need durable collection-file storage. My recommendation remains YuNet + SFace, with SFace's training-data confirmation as an explicit pre-release gate.
Author
Owner

Finished research-only job on job/research-faces. HEAD SHA: 57118d9648582e682f0a0e1997fc8ad9f84bab35 (unchanged from base); worktree is clean. No code or generated files changed, so gates were not run: cargo fmt --check, cargo clippy --all-targets -D warnings, cargo test, bash packages/api-client/check-generated.sh, and bash tests/adversarial/run.sh. There is no gate command output to quote for this research-only job. The research recommendation and evidence are in the preceding comment. Issue remains open.

Finished research-only job on `job/research-faces`. HEAD SHA: `57118d9648582e682f0a0e1997fc8ad9f84bab35` (unchanged from base); worktree is clean. No code or generated files changed, so gates were not run: `cargo fmt --check`, `cargo clippy --all-targets -D warnings`, `cargo test`, `bash packages/api-client/check-generated.sh`, and `bash tests/adversarial/run.sh`. There is no gate command output to quote for this research-only job. The research recommendation and evidence are in the preceding comment. Issue remains open.
Sign in to join this conversation.
No labels
No milestone
No project
No assignees
1 participant
Notifications
Due date
The due date is invalid or out of range. Please use the format "yyyy-mm-dd".

No due date set.

Dependencies

No dependencies set

Reference
kayg/calternal#31
No description provided.