Fine-tuning Liquid for product search
LFM2.5-Embedding-350M, fine-tuned on Japanese Amazon ESCI query–product relevance labels: Exact, Substitute, Complement, and Irrelevant. Training updated the shared query/product encoder; all product embeddings were rebuilt afterward.
Held-out results
500 Japanese queries · 339,058 products · 95% confidence intervals below each score
| Metric | Original | Fine-tuned |
|---|---|---|
| NDCG@10 | 0.716 0.696–0.735 | 0.761 0.742–0.780 |
| Known-Exact Recall@10 | 22.7% 20.3%–25.0% | 27.7% 25.2%–30.3% |
| Known-Exact MRR@10 | 0.464 0.426–0.503 | 0.553 0.515–0.591 |
Training & evaluation
| Data | Amazon ESCI: 6,000 training queries, 500 test queries. |
|---|---|
| Relevance gains | Exact = 1 · Substitute = 0.1 · Complement = 0.01 · Irrelevant = 0. Used for NDCG and weighting training pairs by their gain difference. |
| Training objective | Pairwise logistic loss on cosine similarities, scaled by 20. Update the shared encoder through both query and product inputs. |
| Pair sampling | Sample queries uniformly. Approximately 80% of pairs compare two products with different judged relevance labels for the same query; the lower-rated product is the negative. |
| Random negatives | Approximately 20% pair an Exact product with a uniformly sampled catalog product, excluding known Exact, Substitute, and Complement matches for that query. These sampled negatives have weight 1. |
| Optimization | AdamW · learning rate 1e−5 · weight decay 0.01 · effective batch 16 · 600 updates. 10% warmup, then linear decay. |
| Evaluation | NDCG over judged candidates; Recall and MRR over the full catalog, using known Exact matches. Confidence intervals use 10,000 bootstrap resamples of test queries. |