Credits & analysis modes
The credit system is the consumption unit of the vario Scoring API. This page explains what a credit is, how it is deducted based on the chosen analysis mode, how to check your balance, and what happens when credits run out.
1. What is a credit
A credit is the minimum consumption unit of the API. Each call to POST /api/v1/analyze deducts credits from the tenant's balance before executing the analysis. Credits are not deducted for calls that return an HTTP error (for example, 400, 401, 422).
Credits are prepaid: they are purchased before use. There is no post-consumption billing or automatic debt.
2. Analysis modes and their cost
Each request to /api/v1/analyze accepts a mode field that determines the depth of analysis and the number of credits deducted.
| Mode | Layers executed | Credits per call | Typical speed | Recommended for |
|---|---|---|---|---|
screening | L1 (keyword + embedding) | 1 | ~50 ms | Mass corpus sweep, fast discard of irrelevant documents |
fast | L1 + L2 (light LLM) | 2 | ~200 ms | Moderate-precision classification after initial screening |
full | L1 + L2 + L3 (deep LLM) | 3 | ~2–8 s | High-precision analysis with full reasoning |
screening
The screening mode executes only layer L1: keyword detection and semantic similarity via embeddings. It does not invoke any LLM. It is the most economical mode (1 credit) and the fastest, designed to discard the bulk of a large corpus before applying more costly analysis.
The screening response has a different schema from fast and full: it delivers L1 scores (l1_score, l1_keyword_score, l1_embedding_score) and identified key terms, but does not include L2/L3 fields such as combined_score, should_alert, reasoning, or regulatory_frameworks. See Response object for the complete screening schema.
{
"text": "Hablamos con los del otro grupo y acordamos no movernos del precio.",
"playbook": "price_fixing",
"language": "es",
"mode": "screening"
}
fast
The fast mode executes layers L1 and L2 (light LLM). It offers moderate precision with low latency. It is the recommended mode for the second step of an analysis funnel: it receives documents that exceeded the screening threshold and classifies them before escalating the most relevant ones to full.
{
"text": "Hablamos con los del otro grupo y acordamos no movernos del precio.",
"playbook": "price_fixing",
"language": "es",
"mode": "fast"
}
full
The full mode executes all three layers (L1 + L2 + L3 with deep LLM). It delivers greater precision and more detailed reasoning in the reasoning field. It is the default mode if mode is not specified in the request.
{
"text": "Hablamos con los del otro grupo y acordamos no movernos del precio.",
"playbook": "price_fixing",
"language": "es",
"mode": "full",
"thread_context": [
"Reunión de mañana: definimos rangos de precios para Q3.",
"Los colegas de la competencia nos llamaron para coordinar."
]
}
3. Free trial
Upon registration, each account receives 250 free credits. No credit card is required to access the trial.
With 250 credits you can run up to 250 calls in screening mode, 125 in fast mode, or 83 in full mode. That is enough to evaluate the system with real corpora before committing to a top-up.
Registration is done via the public endpoint POST /api/public/api-signup or from the portal.
curl -X POST https://api.vario.lat/api/public/api-signup \
-H "Content-Type: application/json" \
-d '{"email": "[email protected]", "org_name": "Forensic Firm LATAM"}'
{
"key": "vario_xbGhABC123...WXYZ",
"warning": "Store this key immediately. It cannot be recovered — only revoked and replaced.",
"free_credits": 250,
"call_cost_screening": 1,
"call_cost_fast": 2,
"call_cost_full": 3
}
4. Credit top-ups
When the trial credits run out, you can top up from the portal by choosing the USD amount you want to load. Prices are in USD; Paddle acts as Merchant of Record and determines and remits the applicable taxes based on the buyer's jurisdiction.
For companies with a valid tax ID (CUIT, RFC, CNPJ, RUT, RUC, VAT number), the tax paid at checkout is recoverable as a tax credit.
See the complete checkout reference at Buying credits.
5. Checking your credit balance
Use the GET /api/v1/api-keys/credits endpoint to get the current balance and the estimated remaining calls per mode.
curl https://api.vario.lat/api/v1/api-keys/credits \
-H "Authorization: Bearer vario_YOUR_API_KEY"
Response 200:
{
"balance": 487,
"call_cost_screening": 1,
"call_cost_fast": 2,
"call_cost_full": 3,
"estimated_calls_remaining_screening": 487,
"estimated_calls_remaining_fast": 243,
"estimated_calls_remaining_full": 162
}
| Field | Type | Description |
|---|---|---|
balance | integer | Available credits in the tenant. |
call_cost_screening | integer | Credits deducted per call in screening mode. |
call_cost_fast | integer | Credits deducted per call in fast mode. |
call_cost_full | integer | Credits deducted per call in full mode. |
estimated_calls_remaining_screening | integer | Estimated remaining calls in screening mode. |
estimated_calls_remaining_fast | integer | Estimated remaining calls in fast mode. |
estimated_calls_remaining_full | integer | Estimated remaining calls in full mode. |
6. What happens when credits run out
If the balance reaches zero, the POST /api/v1/analyze endpoint returns 402 INSUFFICIENT_CREDITS:
{
"error": {
"code": "INSUFFICIENT_CREDITS",
"message": "Credit balance is 0. Top up credits to continue."
}
}
No credits are deducted in this case. The request is not processed.
Options to continue:
- From the portal — log in and top up credits in the billing section.
- From the API — check options with
GET /api/public/checkout/credits/previewand initiate checkout withPOST /api/public/checkout/credits. See Buying credits.
7. Best practices for efficient credit usage
The cost per analysis varies 3x between modes. A three-stage funnel significantly reduces consumption without sacrificing precision where it matters.
Recommended pattern — three-stage funnel:
screening— sweep the entire corpus with the L1 layer (keyword + embedding, no LLM). Discard documents with no signal. Cost: 1 credit per message.fast— analyze only the documents that exceeded thescreeningthreshold. Cost: 2 credits per message, applied to a fraction of the corpus.full— process in depth only the documents with a significant signal fromfast. Cost: 3 credits per message, applied to the smallest fraction.
# Example three-stage funnel
for message in corpus:
# Stage 1: mass discard with screening (1 credit)
result_screening = analyze(message, mode="screening")
if result_screening["l1_score"] < 0.3:
continue # weak signal → discard
# Stage 2: classification with fast (2 credits)
result_fast = analyze(message, mode="fast")
if not result_fast["should_alert"]:
continue # below threshold → discard
# Stage 3: deep analysis with full (3 credits)
result_full = analyze(message, mode="full")
# Use result_full for legal review
Additional recommendations:
- Monitor the
balancefield periodically withGET /api/v1/api-keys/credits. Implement alerts when the balance falls below an operational threshold defined by your team. - Use
thread_contextonly infullcalls where thread context adds value. Inscreeningandfastmodes the field is ignored or has limited impact. - Do not use
thread_contextto split messages that exceed the character limit oftext. For long documents, send independent chunks.