Embeddings API
POST /v1/embeddings converts text into numeric vectors, using the same request and response shape as the OpenAI embeddings API. Use it for semantic search, clustering, deduplication, recommendation and retrieval pipelines.
#Quickstart
curl https://api.radienthq.com/v1/embeddings \
-H "Authorization: Bearer $RADIENT_API_KEY" \
-H "Content-Type: application/json" \
-d '{
"model": "auto",
"input": "Embeddings power semantic search, clustering and retrieval."
}'
{
"object": "list",
"data": [
{
"object": "embedding",
"embedding": [ -0.022395002, 0.0020443022, "… 1,536 values total …" ],
"index": 0
}
],
"model": "…",
"usage": { "prompt_tokens": 10, "completion_tokens": 0, "total_tokens": 10, "cost": 0.000001 }
}
model and the vector length are trimmed above: the response reports which embedding model served the request, and that model's vectors. With model: "auto" today the vectors are 1,536 dimensions.
#Request
| Field | Type | Required | Notes |
|---|---|---|---|
model | string | yes in practice | "auto" selects the default embedding model. Any other id is sent to the serving endpoint unchanged; an id the endpoint does not serve fails the request (see Errors). |
input | string or array of strings | yes in practice | The text to embed. An array returns one vector per element, in order (index matches the position). |
encoding_format | string | no | Accepted for OpenAI compatibility; returned vectors are floats. |
user | string | no | Accepted for compatibility. |
There is no server-side range checking: a missing model or input is forwarded as-is and the request fails at the serving step (a 500, below).
#Response
| Field | Meaning |
|---|---|
object | "list" |
data[] | One entry per input: {object: "embedding", embedding: [floats], index} |
model | The embedding model that served the request |
usage | prompt_tokens, completion_tokens, total_tokens (the token count billed) and cost in USD |
usage.cost is the amount deducted from your balance for the call. When a cost is absent, no figure was available; it does not mean the call was free.
#Batch example
curl https://api.radienthq.com/v1/embeddings \
-H "Authorization: Bearer $RADIENT_API_KEY" \
-H "Content-Type: application/json" \
-d '{
"model": "auto",
"input": ["first document", "second document", "third document"]
}'
The response has three data entries, index 0–2. Keep batches to a size the serving model accepts; extremely large batches are refused by the endpoint.
#Errors
| Status | Body | Cause |
|---|---|---|
400 | {"error":"Invalid request format"} | The body is not valid JSON |
401 | {"error":"missing Authorization header"} and the other auth bodies | No or invalid credential (Authentication) |
402 | {"error":"insufficient credits"} | The credit gate refused the call; see the credit gate |
500 | {"error":"Failed to process embeddings request"} | Every serving failure — unknown model, missing input, endpoint errors — is flattened into this one response. The upstream's own message is logged, not returned. |
#Cost
Embeddings are billed per input token at the embedding model's rate, with no markup — the same figure is deducted from your balance that the model charges. A failed call records a zero-usage entry against your account, so failures do not consume credits.
See also: Usage and credits API · API Reference · Rate limits.