Fine-tuning Liquid for product search

LFM2.5-Embedding-350M, fine-tuned on Japanese Amazon ESCI query–product relevance labels: Exact, Substitute, Complement, and Irrelevant. Training updated the shared query/product encoder; all product embeddings were rebuilt afterward.

Held-out results

500 Japanese queries · 339,058 products · 95% confidence intervals below each score

MetricOriginalFine-tuned
NDCG@100.716
0.696–0.735
0.761
0.742–0.780
Known-Exact Recall@1022.7%
20.3%–25.0%
27.7%
25.2%–30.3%
Known-Exact MRR@100.464
0.426–0.503
0.553
0.515–0.591

Training & evaluation

DataAmazon ESCI: 6,000 training queries, 500 test queries.
Relevance gainsExact = 1 · Substitute = 0.1 · Complement = 0.01 · Irrelevant = 0.
Used for NDCG and weighting training pairs by their gain difference.
Training objectivePairwise logistic loss on cosine similarities, scaled by 20. Update the shared encoder through both query and product inputs.
Pair samplingSample queries uniformly. Approximately 80% of pairs compare two products with different judged relevance labels for the same query; the lower-rated product is the negative.
Random negativesApproximately 20% pair an Exact product with a uniformly sampled catalog product, excluding known Exact, Substitute, and Complement matches for that query. These sampled negatives have weight 1.
OptimizationAdamW · learning rate 1e−5 · weight decay 0.01 · effective batch 16 · 600 updates.
10% warmup, then linear decay.
EvaluationNDCG over judged candidates; Recall and MRR over the full catalog, using known Exact matches. Confidence intervals use 10,000 bootstrap resamples of test queries.