Book a call

On-Device AI vs Cloud AI: What We Learned Building a Passport Photo App

Running every check on the phone removed per-photo inference cost and kept user photos local. It also meant fewer models to choose from and a lot of tuning against noisy sensors.

Isometric illustration: a smartphone under a glass dome with a photo card, a cloud far outside

On-device AI trades model choice for privacy, cost and speed

On-device AI means the model runs on the user's phone or in their browser, so the input never has to reach your server. In Photo for Passport, an app that turns an ordinary photo into a compliant passport or ID photo, all 23 compliance checks run on the device using Apple Vision on iOS and MediaPipe in the browser. That gives us zero inference cost per photo, no upload wait, and a privacy promise users can understand: the photo you are checking stays on your phone.

The price is a smaller menu of models and a lot of careful tuning. This is what the trade looked like in practice.

Edge AI vs cloud AI: the short version

  • Cost. Cloud inference is billed per call. On-device inference costs nothing per call once the app is installed, which matters for a product where a user may retake a photo ten times.
  • Latency. No upload and no queue, and the checks keep working offline.
  • Privacy and compliance. A face photo is personal data. If it never leaves the device, most of the data-protection questions never come up.
  • Model choice. The cloud lets you call the largest model available. On the device you work with what fits in memory and what the platform ships, such as Apple Vision, Core ML models and MediaPipe.
  • Updates. A cloud model changes when you deploy. An on-device model changes when users update the app, so thresholds and data need a remote way to be switched off.

For photo compliance, every item on that list favours the device except model choice. That turned out to matter less than expected, because the hard part was never the model.

What runs on the phone

The iOS pipeline is built from Apple Vision requests: face rectangles and landmarks for pose, eyes, mouth and gaze; face capture quality for focus; and person segmentation for the background mask and for finding the shoulders. The checks run in five groups, in a fixed order:

  1. Face and pose: one person, head upright (pass under 5° of roll), shoulders level, facing the camera, eyes open, mouth closed.
  2. Framing: head size against the document's spec, space above the head and below the chin, enough resolution on the face.
  3. Lighting: face brightness, overexposure, even light across both cheeks, harsh highlights, contrast.
  4. Image quality: natural colour and focus.
  5. Background: plain and light, measured in the top corners where hair and shoulders rarely reach.

The crop itself is plain arithmetic from the document's millimetre spec: 258 document types across 61 countries, each with its own size and head-height range. Export sizes are computed from millimetres and DPI, and the app only offers a resolution the source photo can actually support. It never upscales.

The web version follows the same rule. Background removal and face landmarks run in the browser with MediaPipe, and when the user pays, the photo waits in the browser's own storage during the payment redirect instead of being uploaded.

Lenient by design

The most important decision was not about models at all. A check that wrongly tells someone their good photo is bad does more damage than a check that misses a minor issue, because the user loses trust in every other result. So thresholds are deliberately forgiving, and most checks have three states (pass, warn, fail) instead of two. Head tilt under 5° passes and under 11° only warns. Eye openness fails only below a clear threshold and warns in the grey zone.

The counter is honest too. The app shows the number of checks it actually ran on that photo, rather than a fixed marketing number.

Sensor noise is part of the spec

Thresholds that work on one platform can fail on another, because the inputs are noisy in different ways. Two examples from our port to Android:

  • On a perfectly level photo, the Android face-landmark library reported an eye-line tilt of 1.34°. The iOS pipeline straightened anything above 0.7°, which would have rotated good photos for no reason. The Android deadband went up to 3° after measuring the noise.
  • The body-pose estimator was unreliable on tight head-and-shoulders crops, so shoulder positions are taken from the segmentation silhouette instead.

The lesson carries over to any edge AI project: measure the noise floor of each model on real inputs before you copy thresholds across devices.

Where we said no to generative AI

It is tempting to fix a dim or soft photo with a generative model. We tested two cloud enhancement services and removed them. One changed the person's identity enough to be visible, which is unacceptable on an identity document, and several relighting models we liked carried non-commercial licences. Lighting correction is now deterministic image processing with no model involved.

Some authorities forbid edits outright: in our catalogue, 12 documents prohibit face enhancement and 2 prohibit background replacement. For those, a data-driven "safe output" mode turns the edits off.

The hardest bug was not machine learning

Ten audits had confirmed our photo specs were correct. Then a refund request revealed that Germany has, since 2025, accepted passport photos only from certified providers who submit them digitally. The spec was right; the photo still would not be accepted, because our data model had no field for how a photo is submitted.

We added an acceptance layer to the document data and audited the 30 most used documents. 15 of them had a blocker of some kind, covering about two thirds of recent traffic. The app now shows those warnings before the user starts.

That is the broader point about on-device AI. Moving the model to the phone solved cost, latency and privacy. Whether the product is right still depends on data and rules no model can infer.

Photo for Passport is one of the products we build and run ourselves; see the rest in MetaProject Labs. If you are deciding between edge and cloud for an AI feature, our Generative AI & LLM Engineering team can help you weigh it against your data, cost and compliance constraints.

Related: LLM Evaluation: How to Measure LLM Quality

planning a new project
We’ll help you choose the right discovery depth and map a realistic starting plan.
Book a call
Four translucent glowing squares in purple, blue, green, and orange arranged in a row on a white background.

Recent blog posts

Read More Articles
Isometric illustration: a winding road through glass checkpoint frames with a small yellow car
LLM Guardrails in Practice: How Code Checks an AI Trip Planner
View insights

The model proposes, the code decides. How a deterministic engine catches hallucinated distances, impossible days and broken schedules in an AI road-trip planner.