Skip to content
Pastel LabsThe public research side of Pastel-Cloud OÜ

filed under: on-device photo search

Photo search on your phone

This page measures how well a phone handles photo search on its own. It downloads and starts Google's EmbeddingGemma 2 model, turns a photo or a typed query into an embedding (a list of 768 numbers that describes it), and searches 1,000 photos with those numbers.

To time the model, it encodes up to 12 photos from Flickr8k, a public collection of Flickr photos with written captions, bundled with the page. The 1,000 photos it searches come as prepared embeddings. The model runs inside this tab, on the graphics chip through WebGPU when the browser offers it, otherwise on the processor through WebAssembly. This site serves a fixed copy of the public model files, so the page never fetches from another site. Nothing is uploaded, and results stay in this tab until you press Copy results.

The first run downloads 0.28 GB of model files for the smaller q4 version, plus 0.88 GB if the larger fp16 tests are ticked, so use Wi-Fi if you can. A second visit loads them from the browser's cache. Bigger libraries (10,000, 100,000 and 1 million photos) are timed with randomly generated data. The import-queue test reloads the page once on purpose, then carries on by itself.

This page uses our published average photo for EmbeddingGemma 2, built from 18,232 COCO photos (Common Objects in Context, a public photo collection). It is on Hugging Face at pastel-labs/embeddinggemma-2-photo-anchor, commit 12f5978. This site serves its own copy of the file, and the page checks it against the published checksum before searching.

Options
Ready.
Log

Words used on this page

Embedding
A list of 768 numbers the model produces for a photo or a sentence. Similar things get similar lists, which is what makes search work.
q4 and fp16
Two sizes of the same model. q4 stores most of its weights in 4 bits (0.28 GB of files), fp16 in 16 bits (0.88 GB).
q4-wasm
Our edit of the q4 files. One step the browser's WebAssembly runtime does not have (GatherBlockQuantized) is rewritten into standard steps. On our test machine its output matched the original exactly.
WebGPU
A browser interface to the graphics chip. When it is missing, the page reports "no adapter" and skips those tests.
WebAssembly (wasm)
A compact low-level code format that runs on the processor. Threads let it use several processor cores.
Image pieces (tokens)
The model reads a photo as 280 small pieces by default. The 70-piece option tests a smaller, faster input; "cos vs box" shows how close the result stays.
cos vs box
How similar this device's embedding is to our reference embedding for the same photo, made on our own test computer ("box"). It is a cosine score from 0 to 1, and 1.00000 means identical.
Lookup-table search
Photos are stored as one bit per number, and the query keeps full numbers. Per query the page builds 96 small tables, so each photo costs 96 lookups. "Naive" adds all 768 numbers one by one and is the correctness check.
Parity
Whether this device returns exactly the same top 10 photos for each of 100 queries as our reference run. Equal scores go to the photo with the lower number.
Import queue
Photos encoded one at a time from a saved list, the way a real photo import would run. After a reload the page measures how long it takes to load the model from cache and encode the next photo.
Cross-origin isolated
A browser mode this site turns on so timers are precise and memory can be measured. Memory readings are only available in Chrome-based browsers.