Get in touch

9io.ai / Blog

Comparing FashionSigLIP, CLIP and DINOv2 for visual product search

On 60 real cases, FashionSigLIP put the right product in the top five 63% of the time, against 38% for CLIP and 23% for DINOv2. The pipeline and the test.

Key takeaways

  • FashionSigLIP ranked the right product first in 25 of 60 cases, against 14 for CLIP and 7 for DINOv2.
  • At 60 cases, each hit rate carries up to about 12 points of uncertainty either way.
  • Sixty cases separated FashionSigLIP from CLIP, and were too few to separate CLIP from DINOv2.
  • DINOv2 came last at matching products and is used only to pin exact duplicates.
  • A hosted GroundingDINO API ran a batch in 14.5 s, against 21.7 s locally for us.

When a shopper uploads a photo of an outfit, the choice of image embedding has a large effect on whether the right product comes back. On 60 real test cases, Marqo-FashionSigLIP put the right product first 41.7% of the time and in the top five 63.3% of the time. CLIP scored 23.3% and 38.3%. DINOv2 scored 11.7% and 23.3%.

The search runs on a fashion commerce platform we build and run, over a catalogue of about half a million items. Shoppers upload a photo or a video frame of an outfit, and the platform finds matching products. FashionSigLIP is the embedding the platform uses. DINOv2 stays in the pipeline for one narrow job, pinning exact duplicates.

This post covers the pipeline around the embedding, why we think the three models ranked as they did, and what a 60-case test can tell you. That includes confidence intervals, a significance check and the sample sizes you would need to see smaller gaps.

How the three models scored on 60 cases

Model Top-1 hit rate Top-5 hit rate
Marqo-FashionSigLIP 41.7% (25 of 60) 63.3% (38 of 60)
CLIP 23.3% (14 of 60) 38.3% (23 of 60)
DINOv2 11.7% (7 of 60) 23.3% (14 of 60)

A top-1 hit means the right product ranked first. A top-5 hit means it appeared anywhere in the first five results. FashionSigLIP found the right product first almost twice as often as CLIP, 25 cases against 14, and more than three times as often as DINOv2, which managed 7.

Each case is worth 1/60 of the total, about 1.7 points, which is a reminder of how small the test is. The 95% intervals run up to about 12 points either side of each figure. They are in the section on how we measured, along with what the test can and can’t separate.

The pipeline from a shopper’s photo to a product

Every upload goes through the same steps.

  1. Find the person in the photo or frame.
  2. Describe the garments with one Gemini call.
  3. Get a bounding box for each garment from GroundingDINO, through a hosted detection API.
  4. Cut each garment out with SAM 2.
  5. Embed each cutout with Marqo-FashionSigLIP.
  6. Search for the nearest catalogue items in a managed vector database.
  7. Re-rank the candidates in a step with a fixed cap.

DINOv2 sits beside this flow and is used only to catch exact duplicates.

Steps one, three and four narrow what the embedding sees. A photo of an outfit holds several garments, a person and a background, and one embedding of the whole frame blends all of it. A cutout per garment gives one vector per garment, which can be matched against single products in the catalogue.

The description step relies on Gemini’s image understanding. Gemini models1 take images as input and can caption them or answer questions about them without a task-specific model. The re-ranking step works on a capped number of candidates, so its cost per search has a ceiling however large the catalogue grows.

Finding and cutting out each garment

GroundingDINO2 is an open-set object detector. It finds objects described by category names or referring expressions, so the garments it can box aren’t limited to a fixed list of labels. Its authors report 52.5 AP on COCO zero-shot transfer, without training on COCO.

SAM 23 takes a prompt, such as a box, and returns a mask for the object. On images its authors report it is more accurate than the original Segment Anything Model and six times faster. With a box for each garment, SAM 2 can cut each one out of the frame.

We call GroundingDINO through a hosted detection API. The hosted API took 14.5 s per batch, against 21.7 s when we ran the model locally, about a third less time. That comparison depends on the local hardware and the batch size, so time both on your own workload before deciding where a detector should run.

Latency on the text side of the same platform is a separate story, covered in How we halved an AI stylist’s latency, and when hedging stops paying.

Why a fashion-tuned SigLIP beat general CLIP

CLIP4 learned from 400 million image-text pairs collected from the internet, by predicting which caption goes with which image. SigLIP5 kept the idea and changed the loss. It scores each image-text pair on its own with a sigmoid, so it doesn’t need to normalise across the whole batch.

Marqo-FashionSigLIP6 starts from a SigLIP ViT-B/16 model and fine-tunes it on fashion data with a method Marqo calls Generalised Contrastive Learning. The training signal covers categories, style, colours, materials, keywords and fine details as well as text descriptions. The model card lists an Apache 2.0 licence.

Marqo reports gains of up to 57% over FashionCLIP 2.0 on its own benchmark, built from seven public fashion datasets.7 Those figures are vendor-reported. They measure text-to-image, category-to-product and sub-category-to-product retrieval, and none of the tasks is image-to-image, which is what a shopper’s photo needs. They can’t stand in for a test on your own photos. Marqo also publishes Marqo-FashionCLIP8, fine-tuned from a LAION-trained OpenCLIP ViT-B/16. It isn’t part of this comparison.

Our reading of the result is that fashion-specific training teaches the model which differences matter for a product’s identity, such as colour, material and category. Those carry over from a shopper’s photo to a catalogue image. CLIP was trained on general web data with no fashion-specific objective. The 60-case test doesn’t isolate the cause, so treat this as a hypothesis.

Why DINOv2 came last, and the job it keeps

DINOv29 learns visual features from images alone, with no text and no labels. Its training set, LVD-142M, holds 142 million images, picked automatically for being close to images in curated datasets. The paper reports large gains on instance-level retrieval benchmarks, such as the Oxford and Paris landmark sets.

Landmarks are rigid. Garments fold, sit on different bodies and show up in different light, and a shopper’s photo can differ from the catalogue image in pose, lighting and background. Nothing in DINOv2’s training tells it that colour and material matter more than pose or lighting when the goal is to identify a product. That is our reading of why it came last. As with CLIP, the test doesn’t isolate the cause.

Exact duplicates are a different case. When an upload is the same picture as a catalogue image, appearance is all there is to match, and that suits a model trained on appearance alone. That is DINOv2’s one job in the pipeline. It identifies exact duplicates so the matching product can be pinned.

How we measured, and the limits of 60 cases

The test set has 60 real cases. Each pairs an input image with the catalogue product it should find. We ran the same 60 cases through each model and counted a top-1 hit when the right product ranked first, and a top-5 hit when it ranked anywhere in the first five.

A hit rate from 60 cases is a rough number. The table below gives 95% Wilson score intervals, which Brown, Cai and DasGupta10 recommend for small samples in place of the textbook normal approximation.

Model Top-1 (95% interval) Top-5 (95% interval)
Marqo-FashionSigLIP 41.7% (30.1 to 54.3%) 63.3% (50.7 to 74.4%)
CLIP 23.3% (14.4 to 35.4%) 38.3% (27.1 to 51.0%)
DINOv2 11.7% (5.8 to 22.2%) 23.3% (14.4 to 35.4%)

The intervals are 16 to 24 points wide. On a different set of 60 cases, any single figure could plausibly sit 10 points higher or lower.

The large gaps survive that uncertainty. FashionSigLIP leads CLIP by 18.3 points at top-1 and 25 points at top-5. A two-sided two-proportion test gives p ≈ 0.03 and p ≈ 0.006 for those gaps. CLIP’s lead over DINOv2, 11.7 points at top-1 and 15 at top-5, gives p ≈ 0.09 and p ≈ 0.08. So the test shows FashionSigLIP ahead of both, and it can’t separate CLIP from DINOv2 with confidence.

Those tests treat each model’s run as an independent sample. All three models saw the same cases, so a paired test fits better. McNemar’s test compares two models using only the cases where they disagree, and Dietterich11 found it kept false alarms acceptably low when comparing two classifiers on a single test set. When models tend to succeed and fail on the same cases, a paired test usually has more power than the unpaired figures above. We give the unpaired version because anyone can check it from the counts in this post.

Other limits worth stating:

  • The cases come from one catalogue. Another catalogue, with different categories or image styles, may rank the models differently.
  • A hit is the exact product. A near-identical item from the catalogue counts as a miss, though a shopper might be happy with it.
  • The table has one row per model family. Other sizes and versions of CLIP and DINOv2 may score differently.
  • This is a ranking test. It doesn’t measure latency or cost per search.

How many cases your own test needs

Sixty cases were enough here because the gaps were large. Smaller gaps need far more. The standard two-proportion formula, at 80% power and a two-sided 5% significance level, gives these sizes for an unpaired comparison.

Hit rates compared Gap Cases needed per model
63% against 38% 25 points about 62
42% against 23% 19 points about 95
40% against 30% 10 points about 356
63% against 53% 10 points about 382

Our 60 cases sit at the edge for a 25-point gap. To resolve a 10-point gap you need several hundred cases per model, or a paired design on shared cases, which often gets there with fewer.

The scoring takes a few lines. Record the rank of the right product for every case, with None when it doesn’t come back at all, then compute hit rates and intervals.

from math import sqrt

def hits_at(ranks, k):
    """ranks: 1-based rank of the right product for each case, or None if it is missing."""
    return sum(1 for r in ranks if r is not None and r <= k)

def wilson(hits, n, z=1.96):
    """95% Wilson score interval for a hit rate of hits out of n."""
    p = hits / n
    centre = (p + z * z / (2 * n)) / (1 + z * z / n)
    half = z * sqrt(p * (1 - p) / n + z * z / (4 * n * n)) / (1 + z * z / n)
    return centre - half, centre + half

print(wilson(25, 60))  # about (0.30, 0.54), the FashionSigLIP top-1 interval

What we’d do again

  • Build the test set from real inputs, record the right product for each case, and score every candidate model on the same cases.
  • Cut each garment out before embedding. Detection plus segmentation turns one photo into one query per garment.
  • Start from a domain-tuned embedding, then check it on your own query type, since vendor benchmarks may measure different tasks.
  • Report intervals next to hit rates, and keep per-case results so models can be compared with a paired test.
  • Give a pure-vision model like DINOv2 the exact-duplicate job and keep it out of general matching.
  • Time hosted against local inference on your own batches before choosing where a model runs.
  • Cap the re-rank so its cost has a ceiling as the catalogue grows.

  1. Google AI for Developers, “Image understanding”, Gemini API docs, https://ai.google.dev/gemini-api/docs/image-understanding. ↩

  2. Shilong Liu et al., “Grounding DINO: Marrying DINO with Grounded Pre-Training for Open-Set Object Detection”, 2023, https://arxiv.org/abs/2303.05499. ↩

  3. Nikhila Ravi et al., “SAM 2: Segment Anything in Images and Videos”, 2024, https://arxiv.org/abs/2408.00714. ↩

  4. Alec Radford et al., “Learning Transferable Visual Models From Natural Language Supervision”, 2021, https://arxiv.org/abs/2103.00020. ↩

  5. Xiaohua Zhai et al., “Sigmoid Loss for Language Image Pre-Training”, ICCV 2023, https://arxiv.org/abs/2303.15343. ↩

  6. Marqo, Marqo-FashionSigLIP model card, https://huggingface.co/Marqo/marqo-fashionSigLIP. ↩

  7. Marqo, marqo-FashionCLIP repository and benchmark, https://github.com/marqo-ai/marqo-FashionCLIP. ↩

  8. Marqo, Marqo-FashionCLIP model card, https://huggingface.co/Marqo/marqo-fashionCLIP. ↩

  9. Maxime Oquab et al., “DINOv2: Learning Robust Visual Features without Supervision”, 2023, https://arxiv.org/abs/2304.07193. ↩

  10. Lawrence D. Brown, T. Tony Cai and Anirban DasGupta, “Interval Estimation for a Binomial Proportion”, Statistical Science 16(2), 2001, https://doi.org/10.1214/ss/1009213286. ↩

  11. Thomas G. Dietterich, “Approximate Statistical Tests for Comparing Supervised Classification Learning Algorithms”, Neural Computation 10(7), 1998, https://doi.org/10.1162/089976698300017197. ↩

Frequently asked questions

What is the best image embedding model for fashion product search?

In our 60-case test against a fashion catalogue, Marqo-FashionSigLIP led with 41.7% top-1 and 63.3% top-5, ahead of CLIP and DINOv2. Test on your own catalogue and photos, because the intervals at this sample size are wide.

Is DINOv2 good for visual product search?

Not for matching a shopper’s photo to a product in our test, where it came last with 11.7% top-1 and 23.3% top-5. We use it only to spot exact duplicates of catalogue images.

Why segment garments before embedding them?

A photo of an outfit holds several garments, a person and a background. Cutting each garment out with SAM 2 gives one embedding per garment, which can be matched against single products.

How many test cases do I need to compare embedding models?

Sixty cases resolved gaps of 18 to 25 points in our test. Detecting a 10-point gap with 80% power takes roughly 350 to 380 cases per model with an unpaired test. A paired test on shared cases often needs fewer.

What do top-1 and top-5 hit rates mean?

Top-1 is the share of test cases where the correct product ranked first. Top-5 counts a hit when the correct product appears anywhere in the first five results.

Should I run GroundingDINO myself or use a hosted API?

Time both on your own batches. For us a hosted detection API ran a batch in 14.5 s against 21.7 s locally, but the result depends on your hardware and batch size.

Work with us

Building something like this?

9io is a small team of senior engineers with a fractional CTO, and we work by the hour. Send us a note about your product. The reply comes from the person who'd do the work.