docs.vario.lat
This page was machine-translated and is pending human review.

Credits & analysis modes

The credit system is the consumption unit of the vario Scoring API. This page explains what a credit is, how it is deducted based on the chosen analysis mode, how to check your balance, and what happens when credits run out.


1. What is a credit

A credit is the minimum consumption unit of the API. Each call to POST /api/v1/analyze deducts credits from the tenant's balance before executing the analysis. Credits are not deducted for calls that return an HTTP error (for example, 400, 401, 422).

ℹ

Charge on text without signal: if the text submitted in fast or full mode does not pass the minimum L1 signal threshold, the pipeline returns processing_path="l1_discard" in the response (200 OK) but the credit is still charged. To protect credits in corpora with trivial text (greetings, meeting confirmations, blank text), use mode: "screening" first — it is the only mode that always returns L1 scores regardless of the threshold.

Credits are prepaid: they are purchased before use. There is no post-consumption billing or automatic debt.


2. Analysis modes and their cost

Each request to /api/v1/analyze accepts a mode field that determines the depth of analysis and the number of credits deducted.

ModeLayers executedCredits per callTypical speedRecommended for
screeningL1 (keyword + embedding)1~50 msMass corpus sweep, fast discard of irrelevant documents
fastL1 + L2 (light LLM)2~200 msModerate-precision classification after initial screening
fullL1 + L2 + L3 (deep LLM)3~2–8 sHigh-precision analysis with full reasoning

screening

The screening mode executes only layer L1: keyword detection and semantic similarity via embeddings. It does not invoke any LLM. It is the most economical mode (1 credit) and the fastest, designed to discard the bulk of a large corpus before applying more costly analysis.

The screening response has a different schema from fast and full: it delivers L1 scores (l1_score, l1_keyword_score, l1_embedding_score) and identified key terms, but does not include L2/L3 fields such as combined_score, should_alert, reasoning, or regulatory_frameworks. See Response object for the complete screening schema.

{
  "text": "Hablamos con los del otro grupo y acordamos no movernos del precio.",
  "playbook": "price_fixing",
  "language": "es",
  "mode": "screening"
}

fast

The fast mode executes layers L1 and L2 (light LLM). It offers moderate precision with low latency. It is the recommended mode for the second step of an analysis funnel: it receives documents that exceeded the screening threshold and classifies them before escalating the most relevant ones to full.

{
  "text": "Hablamos con los del otro grupo y acordamos no movernos del precio.",
  "playbook": "price_fixing",
  "language": "es",
  "mode": "fast"
}

full

The full mode executes all three layers (L1 + L2 + L3 with deep LLM). It delivers greater precision and more detailed reasoning in the reasoning field. It is the default mode if mode is not specified in the request.

{
  "text": "Hablamos con los del otro grupo y acordamos no movernos del precio.",
  "playbook": "price_fixing",
  "language": "es",
  "mode": "full",
  "thread_context": [
    "Reunión de mañana: definimos rangos de precios para Q3.",
    "Los colegas de la competencia nos llamaron para coordinar."
  ]
}

3. Free trial

Upon registration, each account receives 250 free credits. No credit card is required to access the trial.

With 250 credits you can run up to 250 calls in screening mode, 125 in fast mode, or 83 in full mode. That is enough to evaluate the system with real corpora before committing to a top-up.

Registration is done via the public endpoint POST /api/public/api-signup or from the portal.

curl -X POST https://api.vario.lat/api/public/api-signup \
  -H "Content-Type: application/json" \
  -d '{"email": "[email protected]", "org_name": "Forensic Firm LATAM"}'
{
  "key": "vario_xbGhABC123...WXYZ",
  "warning": "Store this key immediately. It cannot be recovered — only revoked and replaced.",
  "free_credits": 250,
  "call_cost_screening": 1,
  "call_cost_fast": 2,
  "call_cost_full": 3
}

4. Credit top-ups

When the trial credits run out, you can top up from the portal by choosing the USD amount you want to load. Prices are in USD; Paddle acts as Merchant of Record and determines and remits the applicable taxes based on the buyer's jurisdiction.

For companies with a valid tax ID (CUIT, RFC, CNPJ, RUT, RUC, VAT number), the tax paid at checkout is recoverable as a tax credit.

See the complete checkout reference at Buying credits.


5. Checking your credit balance

Use the GET /api/v1/api-keys/credits endpoint to get the current balance and the estimated remaining calls per mode.

curl https://api.vario.lat/api/v1/api-keys/credits \
  -H "Authorization: Bearer vario_YOUR_API_KEY"

Response 200:

{
  "balance": 487,
  "call_cost_screening": 1,
  "call_cost_fast": 2,
  "call_cost_full": 3,
  "estimated_calls_remaining_screening": 487,
  "estimated_calls_remaining_fast": 243,
  "estimated_calls_remaining_full": 162
}
FieldTypeDescription
balanceintegerAvailable credits in the tenant.
call_cost_screeningintegerCredits deducted per call in screening mode.
call_cost_fastintegerCredits deducted per call in fast mode.
call_cost_fullintegerCredits deducted per call in full mode.
estimated_calls_remaining_screeningintegerEstimated remaining calls in screening mode.
estimated_calls_remaining_fastintegerEstimated remaining calls in fast mode.
estimated_calls_remaining_fullintegerEstimated remaining calls in full mode.
ℹ

If the expires_at field is present in your account, it indicates the expiration date of the current credits.


6. What happens when credits run out

If the balance reaches zero, the POST /api/v1/analyze endpoint returns 402 INSUFFICIENT_CREDITS:

{
  "error": {
    "code": "INSUFFICIENT_CREDITS",
    "message": "Credit balance is 0. Top up credits to continue."
  }
}

No credits are deducted in this case. The request is not processed.

Options to continue:

  1. From the portal — log in and top up credits in the billing section.
  2. From the API — check options with GET /api/public/checkout/credits/preview and initiate checkout with POST /api/public/checkout/credits. See Buying credits.

7. Best practices for efficient credit usage

The cost per analysis varies 3x between modes. A three-stage funnel significantly reduces consumption without sacrificing precision where it matters.

Recommended pattern — three-stage funnel:

  1. screening — sweep the entire corpus with the L1 layer (keyword + embedding, no LLM). Discard documents with no signal. Cost: 1 credit per message.
  2. fast — analyze only the documents that exceeded the screening threshold. Cost: 2 credits per message, applied to a fraction of the corpus.
  3. full — process in depth only the documents with a significant signal from fast. Cost: 3 credits per message, applied to the smallest fraction.
# Example three-stage funnel
for message in corpus:
    # Stage 1: mass discard with screening (1 credit)
    result_screening = analyze(message, mode="screening")
    if result_screening["l1_score"] < 0.3:
        continue  # weak signal → discard

    # Stage 2: classification with fast (2 credits)
    result_fast = analyze(message, mode="fast")
    if not result_fast["should_alert"]:
        continue  # below threshold → discard

    # Stage 3: deep analysis with full (3 credits)
    result_full = analyze(message, mode="full")
    # Use result_full for legal review

Additional recommendations:

  • Monitor the balance field periodically with GET /api/v1/api-keys/credits. Implement alerts when the balance falls below an operational threshold defined by your team.
  • Use thread_context only in full calls where thread context adds value. In screening and fast modes the field is ignored or has limited impact.
  • Do not use thread_context to split messages that exceed the character limit of text. For long documents, send independent chunks.

ℹ

vario does not store or retain the analyzed text. Each call to /api/v1/analyze is stateless — the analyzed_at field in the response is the only temporal record of the analysis.

The API output is a risk signal, not a legal determination. It does not substitute a lawyer's review and does not constitute evidence before regulators on its own.