Authentication for open-source models

Is that the real DeepSeek V4 Pro?Kimi K2?Qwen3.5-122B?gpt-oss-120b?

Some providers serve a cheaper model, a compressed copy or a broken setup under the name you picked. authenticated.si checks each call against the real model and tells you what you got.

You're on the list. We'll email you before launch.

Pay per check from prepaid credits. No subscription.

Example check No. AS-0048213 · call 48,213

deepseek/deepseek-v4-pro

via example-host

  • Token count612 / 612
  • Fixed answers10 / 10 match
  • Driftnone in 200 tokens
  • Hard questions17 / 20, ref 17
  • Hidden wordfound at 128k
  • Tool calls10 / 10 valid JSON

Authenticated

96% confidence · 1 credit

Example values. Real checks report their own numbers.

Hover the certificate to see it under UV.

01What goes wrong

Three ways a provider cuts corners to save GPU cost.

Swapped

A different model

You pick DeepSeek V4 Pro and get a smaller, cheaper model billed under the same name.

Quantized

A compressed copy

The right model squeezed to 4-bit. Easy prompts look fine; long code and hard math start to slip.

Misconfigured

A broken setup

The right model with its context window cut from 128k to 32k tokens, or tool calls that come back as broken JSON.

02On the record

It already happens, and people have measured it.

16.6points

How far benchmark scores moved when the only change was which OpenRouter provider served the model.

LessWrong

31of 32

AI safety research codebases that called OpenRouter without pinning a provider.

LessWrong

May 52026

A strict client caught a private-inference gateway serving Qwen3.5-122B in place of DeepSeek-V3.1.

awesome-private-inference

A one-time test isn't enough. Developers testing Kimi providers found some pass at light load and slip when their GPUs are busy. Hacker News thread

03How it works

Each check takes three steps.

  1. Send test prompts

    Your terminal or agent sends a short set of test prompts to the endpoint you already use.

  2. Compare to the real model

    We compare the answers with reference answers we made by running the real model ourselves.

  3. Weigh the signals

    Six signals are combined into one verdict with a confidence level, and the check costs a credit.

04The checks

Six things we look at, the way a legit check works.

A sneaker legit check looks at the stitching, the size tag and the sole, because no single detail settles it. We weigh six signals for the same reason.

SwapS-01

Token count

Every model family splits text into tokens differently, so the bill for the same prompt gives a different model away.

Reference612
Endpoint680
SwapS-02

Fixed answers

With randomness turned off, the real model words the same answer the same way every time.

ReferenceThe capital of Australia is Canberra.

EndpointCanberra is the capital of Australia.

QuantS-03

Drift

A compressed copy starts out matching, then wanders from the reference partway through a long answer.

ReferenceEndpoint
QuantS-04

Hard questions

Tricky math and code prompts are where 4-bit copies lose points first.

Reference17 / 20
Endpoint12 / 20
SetupS-05

Hidden word

We hide a code word at the start of a very long prompt. A provider that quietly cut the context window can't find it.

code wordlast 32k kept
SetupS-06

Tool calls

We ask for a function call. The real setup returns clean JSON; a broken one returns text that your code can't parse.

Valid: {"name":"get_weather","arguments":{"city":"Delhi"}}
Broken: get_weather(city=Delhi
05The verdict

One answer per check, with a confidence level.

Authenticated

Every signal matches the reference for the model you picked.

Likely quantized

Same model family, but answers drift partway through and hard questions slip.

Likely swapped

Token counts and fixed answers point to a different model than the one on your bill.

Misconfigured

The model matches, but the context window is cut short or tool calls break.

06For your terminal and your agents

Run it once from the command line, or let your agent check on a schedule.

TerminalPlanned interface
$ authenticated check deepseek/deepseek-v4-pro \
    --endpoint $PROVIDER_URL --depth quick

  ✓ token count      612 / 612
  ✓ fixed answers    10 / 10
  ✓ drift            none in 200 tokens
  ✓ hidden word      found at 128k

  AUTHENTICATED  96% confidence  1 credit
AgentPlanned interface
# before each batch of work
POST https://api.authenticated.si/v1/checks
{
  "model": "deepseek/deepseek-v4-pro",
  "endpoint": "$PROVIDER_URL",
  "depth": "quick"
}

# 200 OK
{ "verdict": "authenticated", "confidence": 0.96 }

Pay per check. Credits come in prepaid packs.

Three depths. Quick, Standard and Deep.

Any OpenAI-compatible endpoint. Including routers like OpenRouter.

Next

Authenticated providers

We plan to serve open models ourselves. Our own endpoints will run the same checks in public, so you can hold us to the standard we hold everyone else to.

Who's building it

From sneakers to models

Kartikeya Sharma co-founded Dype, a sneaker authentication service, in 2019. authenticated.si brings the same legit check to AI models, built by Bhag Labs, Inc.

07Questions

Plain answers.

Do you see my prompts?

Checks use our own test prompts, sent to the endpoint you name. Your own conversations don't pass through authenticated.si.

Is the verdict a cryptographic proof?

No. It's a weighted verdict with a confidence level, built from six independent signals. Like a legit check, it's strong evidence, and we tell you how strong.

Does my provider have to do anything?

No. We test the endpoint from the outside, the same way you use it, so it works with any provider that offers an OpenAI-compatible API.

Which models can you check?

We start with a short list of popular open models whose reference answers we generate ourselves, and add more as each reference run finishes.

How does pricing work?

You buy a pack of credits and each check spends some, depending on depth. There's no subscription. Prices go live at launch, and people on the early access list hear first.

Early access

Know which model answered before you trust the answer.

Get early access