Skip to content
Radient LogoRadient Documentation

Embeddings API

POST /v1/embeddings converts text into numeric vectors, using the same request and response shape as the OpenAI embeddings API. Use it for semantic search, clustering, deduplication, recommendation and retrieval pipelines.

#Quickstart

bash
curl https://api.radienthq.com/v1/embeddings \
  -H "Authorization: Bearer $RADIENT_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{
    "model": "auto",
    "input": "Embeddings power semantic search, clustering and retrieval."
  }'
json
{
  "object": "list",
  "data": [
    {
      "object": "embedding",
      "embedding": [ -0.022395002, 0.0020443022, "… 1,536 values total …" ],
      "index": 0
    }
  ],
  "model": "…",
  "usage": { "prompt_tokens": 10, "completion_tokens": 0, "total_tokens": 10, "cost": 0.000001 }
}

model and the vector length are trimmed above: the response reports which embedding model served the request, and that model's vectors. With model: "auto" today the vectors are 1,536 dimensions.

#Request

FieldTypeRequiredNotes
modelstringyes in practice"auto" selects the default embedding model. Any other id is sent to the serving endpoint unchanged; an id the endpoint does not serve fails the request (see Errors).
inputstring or array of stringsyes in practiceThe text to embed. An array returns one vector per element, in order (index matches the position).
encoding_formatstringnoAccepted for OpenAI compatibility; returned vectors are floats.
userstringnoAccepted for compatibility.

There is no server-side range checking: a missing model or input is forwarded as-is and the request fails at the serving step (a 500, below).

#Response

FieldMeaning
object"list"
data[]One entry per input: {object: "embedding", embedding: [floats], index}
modelThe embedding model that served the request
usageprompt_tokens, completion_tokens, total_tokens (the token count billed) and cost in USD

usage.cost is the amount deducted from your balance for the call. When a cost is absent, no figure was available; it does not mean the call was free.

#Batch example

bash
curl https://api.radienthq.com/v1/embeddings \
  -H "Authorization: Bearer $RADIENT_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{
    "model": "auto",
    "input": ["first document", "second document", "third document"]
  }'

The response has three data entries, index 0–2. Keep batches to a size the serving model accepts; extremely large batches are refused by the endpoint.

#Errors

StatusBodyCause
400{"error":"Invalid request format"}The body is not valid JSON
401{"error":"missing Authorization header"} and the other auth bodiesNo or invalid credential (Authentication)
402{"error":"insufficient credits"}The credit gate refused the call; see the credit gate
500{"error":"Failed to process embeddings request"}Every serving failure — unknown model, missing input, endpoint errors — is flattened into this one response. The upstream's own message is logged, not returned.

#Cost

Embeddings are billed per input token at the embedding model's rate, with no markup — the same figure is deducted from your balance that the model charges. A failed call records a zero-usage entry against your account, so failures do not consume credits.

See also: Usage and credits API · API Reference · Rate limits.