- Task: 3.2
- Evaluated: 2026-09-05
- Decision basis: pricing-and-posture
- Chosen:
perplexity/pplx-embed-v1-0.6b - Registry leg:
embedding— bound - Probe spend: $0.28906573 over 903 cost-reported requests
- Record digest:
86122d349529f9682e6f0f8b38c1939489c29d1e4203b24a97d7347e8fcdd219
What was measured, and why a catalogue was not enough#
OpenRouter's default model listing omits embedding and rerank models entirely,
so the candidate set was ASKED for (/models?output_modalities=embeddings,
/models?output_modalities=rerank) rather than guessed: 37 embedding candidates
and 7 rerankers. Every price below is usage.cost returned by a completed call,
never pricing.prompt from the catalogue — one reranker in this table lists its
prompt price as 0 and bills a real amount.
The corpus is exact: 56,970 chunks, 36,147,926 characters, longest chunk 1205 characters (319 tokens at the widest tokenizer measured). Token volume is projected by a ratio estimator over disjoint calibration batches, with its interval.
Admitted candidates, cheapest first#
| Slug | Dims | $/1M | Index build | Per query | Query p50 | Endpoints | Caveats |
|---|---|---|---|---|---|---|---|
perplexity/pplx-embed-v1-0.6b |
1024 | $0.0040 | $0.0359 ($0.0354–$0.0364) | $3.07e-8 | 203.5 ms | 1 | single-endpoint-no-failover |
baai/bge-base-en-v1.5 |
768 | $0.0050 | $0.0471 ($0.0462–$0.0479) | $5.17e-8 | 309.4 ms | 1 | single-endpoint-no-failover, quantization-undeclared |
thenlper/gte-large |
1024 | $0.0100 | $0.0941 ($0.0924–$0.0959) | $1.03e-7 | 320.6 ms | 1 | single-endpoint-no-failover, quantization-undeclared |
baai/bge-large-en-v1.5 |
1024 | $0.0100 | $0.0941 ($0.0924–$0.0959) | $1.03e-7 | 325.3 ms | 1 | single-endpoint-no-failover, quantization-undeclared |
baai/bge-m3 |
1024 | $0.0100 | $0.1066 ($0.1049–$0.1083) | $9.33e-8 | 338.9 ms | 2 | — |
qwen/qwen3-embedding-8b |
4096 | $0.0100 | $0.0903 ($0.0890–$0.0916) | $8.67e-8 | 608.2 ms | 3 | cross-endpoint-reordering, requires-documented-query-protocol |
voyageai/voyage-4-lite |
1024 | $0.0200 | $0.1795 ($0.1769–$0.1821) | $1.53e-7 | 244.4 ms | 1 | single-endpoint-no-failover, zero-data-retention-unavailable, quantization-undeclared |
openai/text-embedding-3-small |
1536 | $0.0200 | $0.1749 ($0.1722–$0.1776) | $1.40e-7 | 449.1 ms | 2 | quantization-undeclared |
qwen/qwen3-embedding-4b |
2560 | $0.0200 | $0.1807 ($0.1780–$0.1833) | $1.73e-7 | 504.6 ms | 1 | single-endpoint-no-failover, quantization-undeclared, requires-documented-query-protocol |
perplexity/pplx-embed-v1-4b |
2560 | $0.0300 | $0.2693 ($0.2654–$0.2731) | $2.30e-7 | 209.1 ms | 1 | single-endpoint-no-failover |
voyageai/voyage-4 |
1024 | $0.0600 | $0.5385 ($0.5308–$0.5463) | $4.60e-7 | 212.3 ms | 1 | single-endpoint-no-failover, zero-data-retention-unavailable, quantization-undeclared |
mistralai/mistral-embed-2312 |
1024 | $0.1000 | $1.0910 ($1.0769–$1.1050) | $1.07e-6 | 160.3 ms | 3 | quantization-undeclared |
openai/text-embedding-ada-002 |
1536 | $0.1000 | $0.8746 ($0.8612–$0.8880) | $7.00e-7 | 471.5 ms | 1 | single-endpoint-no-failover, zero-data-retention-unavailable, quantization-undeclared |
voyageai/voyage-multimodal-3.5 |
1024 | $0.1200 | $1.0770 ($1.0616–$1.0925) | $9.20e-7 | 269.7 ms | 1 | dimensions-silently-ignored, single-endpoint-no-failover, zero-data-retention-unavailable, quantization-undeclared |
voyageai/voyage-4-large |
1024 | $0.1200 | $1.0770 ($1.0616–$1.0925) | $9.20e-7 | 240.8 ms | 1 | single-endpoint-no-failover, zero-data-retention-unavailable, quantization-undeclared |
openai/text-embedding-3-large |
3072 | $0.1300 | $1.1370 ($1.1195–$1.1544) | $9.10e-7 | 513.8 ms | 2 | cross-endpoint-reordering, quantization-undeclared |
google/gemini-embedding-001 |
3072 | $0.1500 | $1.4047 ($1.3785–$1.4309) | $1.15e-6 | 361.6 ms | 2 | quantization-undeclared |
mistralai/codestral-embed-2505 |
1536 | $0.1500 | $1.4203 ($1.4006–$1.4400) | $1.50e-6 | 201.4 ms | 3 | quantization-undeclared |
google/gemini-embedding-2 |
3072 | $0.2000 | $1.8729 ($1.8380–$1.9079) | $1.53e-6 | 413.3 ms | 4 | quantization-undeclared |
google/gemini-embedding-2-preview |
3072 | $0.2000 | $1.8729 ($1.8380–$1.9079) | $1.53e-6 | 335.3 ms | 1 | single-endpoint-no-failover, zero-data-retention-unavailable, quantization-undeclared |
20 of 37 candidates were admitted; 17 were refused.
Refused, with the measurement that refused them#
| Slug | Refusals | The measurement |
|---|---|---|
liquid/lfm-2.5-embedding-350m:free |
trains-on-input | truncation: refused-oversized-input; no endpoint accepts data_collection: deny |
voyageai/voyage-code-4 |
identity-floor-failed, uncalibrated | plain/nav-admin-repo-map first relevant at rank 6 |
nvidia/nemotron-3-embed-1b:free |
trains-on-input | no endpoint accepts data_collection: deny |
google/gemini-embedding-2:batch |
batch-only-route, unpriced, no-dimensions-observed, trains-on-input, identity-floor-failed, uncalibrated | serves only through the batch API; no cost-reported call |
nvidia/llama-nemotron-embed-vl-1b-v2:free |
trains-on-input | no endpoint accepts data_collection: deny |
thenlper/gte-base |
silently-truncates, uncalibrated | truncation: token-count-plateaus at 512 tokens |
intfloat/e5-large-v2 |
silently-truncates, uncalibrated | truncation: token-count-plateaus at 512 tokens |
intfloat/e5-base-v2 |
silently-truncates, uncalibrated | truncation: token-count-plateaus at 512 tokens |
intfloat/multilingual-e5-large |
silently-truncates, uncalibrated | truncation: token-count-plateaus at 512 tokens |
sentence-transformers/paraphrase-minilm-l6-v2 |
silently-truncates, context-below-corpus-maximum, uncalibrated | truncation: token-count-plateaus at 128 tokens |
sentence-transformers/all-minilm-l12-v2 |
silently-truncates, identity-floor-failed, context-below-corpus-maximum, uncalibrated | truncation: token-count-plateaus at 128 tokens; plain/id-admin-adr-0076 first relevant at rank 8 |
sentence-transformers/multi-qa-mpnet-base-dot-v1 |
silently-truncates, uncalibrated | truncation: token-count-plateaus at 512 tokens |
sentence-transformers/all-mpnet-base-v2 |
silently-truncates, uncalibrated | truncation: token-count-plateaus at 384 tokens |
sentence-transformers/all-minilm-l6-v2 |
silently-truncates, context-below-corpus-maximum, uncalibrated | truncation: token-count-plateaus at 256 tokens |
openai/text-embedding-ada-002:batch |
batch-only-route, unpriced, no-dimensions-observed, trains-on-input, identity-floor-failed, uncalibrated | serves only through the batch API; no cost-reported call |
openai/text-embedding-3-large:batch |
batch-only-route, unpriced, no-dimensions-observed, trains-on-input, identity-floor-failed, uncalibrated | serves only through the batch API; no cost-reported call |
openai/text-embedding-3-small:batch |
batch-only-route, unpriced, no-dimensions-observed, trains-on-input, identity-floor-failed, uncalibrated | serves only through the batch API; no cost-reported call |
The identity floor#
Not a relevance measurement — task 3.5 owns those. This is the bar the cost rule implies: the rule binds "the cheapest model that CAN DO THE JOB", so a price ranking with no capability bar underneath it would rank models that cannot retrieve at all. Each candidate must put a page named by its own identifier above 200 chunks drawn from pages the label does not point at, using only the corpus-identifier labels task 3.1 settled without any ranker.
Several of these models are asymmetric by design and publish a query-side prefix. A model called without the one it documents is a model called wrong, so each is measured under every protocol its family documents AND under plain text as the control.
| Slug | Protocol | Result |
|---|---|---|
liquid/lfm-2.5-embedding-350m:free |
plain | plain/nav-admin-repo-map: pass @1; plain/id-admin-adr-0076: pass @2 |
voyageai/voyage-code-4 |
none cleared | plain/nav-admin-repo-map: FAIL @6; plain/id-admin-adr-0076: pass @1 |
voyageai/voyage-multimodal-3.5 |
plain | plain/nav-admin-repo-map: pass @1; plain/id-admin-adr-0076: pass @1 |
voyageai/voyage-4-lite |
plain | plain/nav-admin-repo-map: pass @1; plain/id-admin-adr-0076: pass @1 |
voyageai/voyage-4 |
plain | plain/nav-admin-repo-map: pass @1; plain/id-admin-adr-0076: pass @1 |
voyageai/voyage-4-large |
plain | plain/nav-admin-repo-map: pass @1; plain/id-admin-adr-0076: pass @1 |
nvidia/nemotron-3-embed-1b:free |
plain | plain/nav-admin-repo-map: pass @1; plain/id-admin-adr-0076: pass @1 |
google/gemini-embedding-2 |
plain | plain/nav-admin-repo-map: pass @1; plain/id-admin-adr-0076: pass @1 |
google/gemini-embedding-2-preview |
plain | plain/nav-admin-repo-map: pass @1; plain/id-admin-adr-0076: pass @1 |
perplexity/pplx-embed-v1-4b |
plain | plain/nav-admin-repo-map: pass @1; plain/id-admin-adr-0076: pass @1 |
perplexity/pplx-embed-v1-0.6b |
plain | plain/nav-admin-repo-map: pass @1; plain/id-admin-adr-0076: pass @3 |
nvidia/llama-nemotron-embed-vl-1b-v2:free |
plain | plain/nav-admin-repo-map: pass @1; plain/id-admin-adr-0076: pass @1 |
thenlper/gte-base |
plain | plain/nav-admin-repo-map: pass @1; plain/id-admin-adr-0076: pass @1 |
thenlper/gte-large |
plain | plain/nav-admin-repo-map: pass @1; plain/id-admin-adr-0076: pass @1 |
intfloat/e5-large-v2 |
plain | plain/nav-admin-repo-map: pass @1; plain/id-admin-adr-0076: pass @1; e5-query-passage/nav-admin-repo-map: pass @1; e5-query-passage/id-admin-adr-0076: pass @1 |
intfloat/e5-base-v2 |
plain | plain/nav-admin-repo-map: pass @1; plain/id-admin-adr-0076: pass @1; e5-query-passage/nav-admin-repo-map: pass @1; e5-query-passage/id-admin-adr-0076: pass @1 |
intfloat/multilingual-e5-large |
plain | plain/nav-admin-repo-map: pass @1; plain/id-admin-adr-0076: pass @2; e5-query-passage/nav-admin-repo-map: pass @1; e5-query-passage/id-admin-adr-0076: pass @1 |
sentence-transformers/paraphrase-minilm-l6-v2 |
plain | plain/nav-admin-repo-map: pass @1; plain/id-admin-adr-0076: pass @3 |
sentence-transformers/all-minilm-l12-v2 |
none cleared | plain/nav-admin-repo-map: pass @1; plain/id-admin-adr-0076: FAIL @8 |
baai/bge-base-en-v1.5 |
plain | plain/nav-admin-repo-map: pass @1; plain/id-admin-adr-0076: pass @1; bge-en-instruction/nav-admin-repo-map: pass @1; bge-en-instruction/id-admin-adr-0076: pass @1 |
sentence-transformers/multi-qa-mpnet-base-dot-v1 |
plain | plain/nav-admin-repo-map: pass @1; plain/id-admin-adr-0076: pass @4 |
baai/bge-large-en-v1.5 |
plain | plain/nav-admin-repo-map: pass @1; plain/id-admin-adr-0076: pass @1; bge-en-instruction/nav-admin-repo-map: pass @1; bge-en-instruction/id-admin-adr-0076: pass @1 |
baai/bge-m3 |
plain | plain/nav-admin-repo-map: pass @1; plain/id-admin-adr-0076: pass @1 |
sentence-transformers/all-mpnet-base-v2 |
plain | plain/nav-admin-repo-map: pass @1; plain/id-admin-adr-0076: pass @2 |
sentence-transformers/all-minilm-l6-v2 |
plain | plain/nav-admin-repo-map: pass @1; plain/id-admin-adr-0076: pass @3 |
mistralai/mistral-embed-2312 |
plain | plain/nav-admin-repo-map: pass @1; plain/id-admin-adr-0076: pass @1 |
google/gemini-embedding-001 |
plain | plain/nav-admin-repo-map: pass @1; plain/id-admin-adr-0076: pass @1 |
openai/text-embedding-ada-002 |
plain | plain/nav-admin-repo-map: pass @1; plain/id-admin-adr-0076: pass @3 |
mistralai/codestral-embed-2505 |
plain | plain/nav-admin-repo-map: pass @1; plain/id-admin-adr-0076: pass @1 |
openai/text-embedding-3-large |
plain | plain/nav-admin-repo-map: pass @1; plain/id-admin-adr-0076: pass @2 |
openai/text-embedding-3-small |
plain | plain/nav-admin-repo-map: pass @1; plain/id-admin-adr-0076: pass @2 |
qwen/qwen3-embedding-8b |
qwen3-instruct | plain/nav-admin-repo-map: pass @1; plain/id-admin-adr-0076: FAIL @7; qwen3-instruct/nav-admin-repo-map: pass @1; qwen3-instruct/id-admin-adr-0076: pass @2 |
qwen/qwen3-embedding-4b |
qwen3-instruct | plain/nav-admin-repo-map: pass @1; plain/id-admin-adr-0076: FAIL @16; qwen3-instruct/nav-admin-repo-map: pass @1; qwen3-instruct/id-admin-adr-0076: pass @1 |
The floor uses 2 corpus-identifier queries against 200 distractor chunks. Its
protocols: plain (no prefix; the control every candidate is measured under);
e5-query-passage (intfloat E5 family model cards: "query: " and "passage: "
prefixes); bge-en-instruction (BAAI bge-*-en-v1.5 model cards: a query-side
retrieval instruction); qwen3-instruct (Qwen3-Embedding model cards: an
instruction-templated query side).
Provider data posture#
Posture is measured by asking the router for it. A request that routes had an endpoint meeting the filter; one that 404s had none, which the router could only know by checking every endpoint.
| Slug | Baseline | Zero data retention | Training denied |
|---|---|---|---|
liquid/lfm-2.5-embedding-350m:free |
routed | refused | refused |
voyageai/voyage-code-4 |
routed | refused | routed |
voyageai/voyage-multimodal-3.5 |
routed | refused | routed |
voyageai/voyage-4-lite |
routed | refused | routed |
voyageai/voyage-4 |
routed | refused | routed |
voyageai/voyage-4-large |
routed | refused | routed |
nvidia/nemotron-3-embed-1b:free |
routed | refused | refused |
google/gemini-embedding-2 |
routed | routed | routed |
google/gemini-embedding-2-preview |
routed | refused | routed |
perplexity/pplx-embed-v1-4b |
routed | routed | routed |
perplexity/pplx-embed-v1-0.6b |
routed | routed | routed |
nvidia/llama-nemotron-embed-vl-1b-v2:free |
routed | refused | refused |
thenlper/gte-base |
routed | routed | routed |
thenlper/gte-large |
routed | routed | routed |
intfloat/e5-large-v2 |
routed | routed | routed |
intfloat/e5-base-v2 |
routed | routed | routed |
intfloat/multilingual-e5-large |
routed | routed | routed |
sentence-transformers/paraphrase-minilm-l6-v2 |
routed | routed | routed |
sentence-transformers/all-minilm-l12-v2 |
routed | routed | routed |
baai/bge-base-en-v1.5 |
routed | routed | routed |
sentence-transformers/multi-qa-mpnet-base-dot-v1 |
routed | routed | routed |
baai/bge-large-en-v1.5 |
routed | routed | routed |
baai/bge-m3 |
routed | routed | routed |
sentence-transformers/all-mpnet-base-v2 |
routed | routed | routed |
sentence-transformers/all-minilm-l6-v2 |
routed | routed | routed |
mistralai/mistral-embed-2312 |
routed | routed | routed |
google/gemini-embedding-001 |
routed | routed | routed |
openai/text-embedding-ada-002 |
routed | refused | routed |
mistralai/codestral-embed-2505 |
routed | routed | routed |
openai/text-embedding-3-large |
routed | routed | routed |
openai/text-embedding-3-small |
routed | routed | routed |
qwen/qwen3-embedding-8b |
routed | routed | routed |
qwen/qwen3-embedding-4b |
routed | routed | routed |
Reranking, priced against what it would be added to#
| Slug | Served by | Billing unit | $/query | Latency | × the embedding query cost |
|---|---|---|---|---|---|
qwen/qwen3-reranker-8b |
Fireworks | unknown | $0.0022886 | 567.2 ms | 74628.26× |
voyageai/rerank-2.5-lite |
VoyageAI by MongoDB | unknown | $0.00015334 | 389.8 ms | 5000.22× |
voyageai/rerank-2.5 |
VoyageAI by MongoDB | unknown | $0.00038335 | 358.5 ms | 12500.54× |
nvidia/llama-nemotron-rerank-vl-1b-v2:free |
Nvidia | unknown | $0 | 729.3 ms | 0× |
cohere/rerank-4-pro |
Cohere | search-unit | $0.0025 | 521 ms | 81521.74× |
cohere/rerank-4-fast |
Cohere | search-unit | $0.002 | 366.5 ms | 65217.39× |
cohere/rerank-v3.5 |
Cohere | search-unit | $0.001 | 279.6 ms | 32608.7× |
deferred to task 3.4 — priced here and not adopted; the measured per-query cost is the reason it needs a quality argument before it is added, not a footnote.
One row bills nothing, and a zero there is a price, not a posture: every :free
route in the embedding field was refused because no endpoint of it accepts
data_collection: deny, and the refusal named "Free model training". A free
reranker earns the same check before it is adopted.
The binding#
perplexity/pplx-embed-v1-0.6b is bound because it is the cheapest candidate
that cleared every measured gate at 10,000 queries/month — $0.0040 per million
tokens, 1024 dimensions, served by Perplexity. The runner-up is
baai/bge-base-en-v1.5, $0.3352 per month dearer at that volume.
The binding is the slug, the endpoint set AND the query protocol plain — an
index built under one convention and queried under another is two systems, not
one.
An endpoint pin is still applied: this slug has a single endpoint, so no agreement between endpoints was measured and none is claimed; the pin names the endpoint that was.
Rollback#
- Slug override:
OSHUN_ASSISTANT_EMBEDDING_MODEL(registry supplies the default underneath it) - Endpoint override:
OPENROUTER_PROVIDER_ONLY - Degraded mode:
lexical-only— apps/oshun/bff/src/assistant/docs-search.ts - Rolling back invalidates stored vectors: true
What task 3.5 must measure#
Task 3.2 binds on price alone, and a price-only pin handed to 3.5 would leave it measuring one arm chosen by cost. The shortlist below is the admitted price/width frontier — for every admitted candidate NOT on it, some admitted candidate is at least as cheap and at least as wide. Dimensions are a capacity fact here, not a claim about retrieval.
perplexity/pplx-embed-v1-0.6bqwen/qwen3-embedding-8b
What re-opens this binding#
- task 3.5 measures relevance and a different candidate wins on a paired test
- task 3.7 declines to promote dense retrieval at all, which retires this leg
- the bound endpoint set changes, because vectors are only comparable within one endpoint
- a re-priced route moves the chosen candidate out of first place at the recorded volume
- the corpus chunker changes, because the calibrated tokens-per-character no longer applies
- the query-side protocol changes, because an index built under one convention cannot be queried under another
- the bound slug gains or loses an endpoint whose quantization differs from the measured one
- the vendor re-points a versioned slug, because a vendor-versioned name is a pin only while the vendor keeps it one
Honest limits#
- This binds on PRICE and POSTURE. No relevance measurement exists yet: task 3.5 owns Recall@k and nDCG over the full set with a pre-registered paired test, and task 3.7 owns promotion. A candidate that clears the identity floor has been shown able to retrieve, not shown to retrieve well.
- The identity floor uses only the relevance set's corpus-identifier labels, because those are the ones that settle without any ranker. Four of the nine query classes still carry no labels at all (task 3.1), so no floor exists for conceptual, multi-hop, recency or contradictory queries.
- Corpus token counts are a ratio estimator over disjoint sampled batches, not an exact count. The corpus CHARACTER count is exact; buying an exact token count would mean embedding all 36 million characters for every candidate.
- Monthly figures are scenarios. This repository records no
search_docscall rate, so the per-query and per-index prices are the measured facts and the monthly totals are arithmetic over a volume a reader chooses. - Latency is wall-clock from this host on this date, over a residential-grade path to one region. It is a floor for planning, not a service-level measurement.
- Posture is measured as ROUTING behaviour — whether an endpoint meeting a filter exists — not as an audit of any provider's actual practice.
- Nothing is indexed. Task 3.3 builds the index, and until it does the bound
slug has no production consumer;
search_docsremains lexical BM25F. - Prices are quoted for text as the corpus stores it. A binding whose query protocol adds a prefix costs a few tokens more per query and, where the protocol also prefixes documents, roughly two per cent more per index build; neither shifts the ranking between candidates at these magnitudes, and the protocol is recorded so the difference is recomputable.
- Each price is the single
usage.costfigure the route returned for one call, retained verbatim in the probe record beside its token count. This record therefore distinguishes no cached from uncached rate, and nothing here should be read as a claim that one exists or does not. - The identity floor is measured under the protocols this tool knows about. A model documenting a calling convention nobody here recognised is measured plain and would be under-measured exactly as the instruction-tuned models were before their protocols were added.