Ainzy API 使用指南Getting Started
使用指南Getting Started

把 Ainzy 接进你的工具和代码

Connect Ainzy to your tools and code

接口与 OpenAI、Anthropic 官方格式一致。现有代码和工具只需要换 Base URL 和 key,别的不用改。

The API follows the OpenAI and Anthropic formats exactly. Existing code and tools only need a new base URL and key. Nothing else changes.

三步开始

Three steps

注册、充值、建令牌,然后就可以调用了。

Register, add credit, create a key. Then start calling.

01

注册账号

Register

用邮箱注册,填入邮件里的验证码即可登录。已有账号直接登录。

Sign up with your email and enter the code we send you. Already have an account? Sign in.

02

充值

Add credit

进入钱包选择金额,支付宝付款。余额以美元计,按 ¥7 = $1 结算。

Open the wallet and pick an amount and pay with Alipay. Balance is in US dollars; Alipay settles at ¥7 = $1.

03

创建令牌

Create a key

在令牌页新建令牌,选好分组后保存,复制以 sk- 开头的 key。分组决定这把 key 能调哪些模型、按什么价格计费。

On the keys page create a key, pick a group, save, and copy the key starting with sk-. The group decides which models the key can call and at what price.

一把 key 只属于一个分组,分组决定它能用的模型。

A key belongs to exactly one group, and the group decides which models it can use.

接入地址

Endpoints

同一把 key、同一个域名,三种协议都可以用,都支持流式输出。

Same key, same host, three protocols. Streaming works on all of them.

协议Protocol地址URL适用Use with
OpenAI Chathttps://ainzy.net/v1/chat/completionsOpenAI SDK、Cursor、Cline、Roo Code、绝大多数第三方应用OpenAI SDKs, Cursor, Cline, Roo Code, most third-party apps
OpenAI Responseshttps://ainzy.net/v1/responsesCodex CLI、使用 Responses API 的新版 SDKCodex CLI and SDKs that use the Responses API
Anthropic Messageshttps://ainzy.net/v1/messagesClaude Code、Anthropic SDKClaude Code, Anthropic SDKs
模型列表Model listhttps://ainzy.net/v1/models返回当前这把 key 所在分组能用的模型名Returns the model names available to this key's group

Base URL 到底填哪个

Which base URL to enter

  • OpenAI 系工具和 SDK(base_url / OPENAI_BASE_URL):填 https://ainzy.net/v1
  • Claude Code 和 Anthropic SDK(ANTHROPIC_BASE_URL):填 https://ainzy.net,不带 /v1
  • 鉴权统一用请求头 Authorization: Bearer sk-你的key;Anthropic 协议也接受 x-api-key
  • OpenAI-style tools and SDKs (base_url / OPENAI_BASE_URL): https://ainzy.net/v1
  • Claude Code and Anthropic SDKs (ANTHROPIC_BASE_URL): https://ainzy.net, without /v1
  • Authenticate with the header Authorization: Bearer sk-YOUR_KEY; the Anthropic protocol also accepts x-api-key

不确定模型名时,先请求一次 /v1/models,列表里有的就能用。模型名区分大小写。

Not sure about a model name? Call /v1/models first. Anything in the list works. Model names are case-sensitive.

分组与模型

Groups & models

创建令牌时选择分组。价格以价格页为准,页面会按你的分组实时显示每百万 token 的单价。

Pick a group when creating a key. Prices live on the pricing page, which shows the per-million-token rate for your group.

分组Group模型Models说明Notes
国模高缓存组Chinese models, 30% of list, high cache
chinese-model-3
deepseek-v4-flash deepseek-v4-flash-0731 deepseek-v4.1-flash deepseek-v4-pro
glm-5.3 glm-5.2 glm-5.3-flash
kimi-k3
DeepSeek、智谱 GLM、Kimi。按官方牌价 3 折计费,缓存命中另按缓存价DeepSeek, Zhipu GLM and Kimi at 30% of official list price. Cache hits are billed at the cache rate
国模组Chinese models, 50% of list
chinese-model-flash
deepseek-v4.1-flash DeepSeek V4.1 Flash,支持 reasoning_effortDeepSeek V4.1 Flash, supports reasoning_effort
阿里云国模组Aliyun Chinese models, 45% of list
aliyun-chinese-model
qwen3.8-max qwen3.8-max-0902 qwen3.8-flash qwen3.7-max qwen3.7-plus qwen3.7-flash qwen3.6-flash
deepseek-v4-pro deepseek-v4-pro-0813 deepseek-v4-flash deepseek-v4-flash-0731 deepseek-v4.1-flash
kimi-k3 kimi-k2.7-code kimi-k2.6 kimi-k2.5
glm-5.1 glm-5.2-fast-preview
doubao-seed-2-1-pro doubao-seed-2-1-turbo
阿里云百炼直连线路,通义千问、豆包、DeepSeek、Kimi、GLM 全按官方牌价 4.5 折计费;与上面两组同名的模型线路独立、互不混用Direct Aliyun Bailian route: Qwen, Doubao, DeepSeek, Kimi and GLM at 45% of official list price. Models sharing a name with the groups above run on a separate, non-mixed route

工具接入

Tools

常用编程工具的配置方法。图形化切换请用 CC Switch,它会替你管理这些配置文件,并且能一键拉取你这把 key 能用的模型。

Configs for common coding tools. For a graphical switcher use CC Switch. It manages these config files for you and can fetch the models your key can use.

用 CC Switch 或任何工具导入国产模型的 key 之后,模型名也必须一起改掉。Claude Code 默认请求 claude- 开头的模型,国产模型分组里没有这个名字,不改就会报 model_not_found。在 CC Switch 里点一次「拉取模型」,或手动填 glm-5.3、kimi-k3。

After importing a Chinese-models key into CC Switch or any other tool, change the model name as well. Claude Code requests claude- models by default; those names do not exist in the Chinese-models groups, so you will get model_not_found. Click "fetch models" once in CC Switch, or type glm-5.3 or kimi-k3 manually.

Claude Code

令牌选「国模高缓存组」,把下面这段加进 ~/.claude/settings.json(Windows 在 %USERPROFILE%\.claude\settings.json):

Create a key in the chinese-model-3 group and add this to ~/.claude/settings.json (on Windows: %USERPROFILE%\.claude\settings.json):

{
  "env": {
    "ANTHROPIC_BASE_URL": "https://ainzy.net",
    "ANTHROPIC_AUTH_TOKEN": "sk-你的key",
    "ANTHROPIC_MODEL": "glm-5.3",
    "ANTHROPIC_SMALL_FAST_MODEL": "glm-5.3-flash"
  }
}
{
  "env": {
    "ANTHROPIC_BASE_URL": "https://ainzy.net",
    "ANTHROPIC_AUTH_TOKEN": "sk-YOUR_KEY",
    "ANTHROPIC_MODEL": "glm-5.3",
    "ANTHROPIC_SMALL_FAST_MODEL": "glm-5.3-flash"
  }
}

也可以直接设环境变量后运行 claude。模型名可换成同分组里的 kimi-k3、deepseek-v4-pro 等。

Setting the same environment variables before running claude also works. Swap the model for any other in the group, such as kimi-k3 or deepseek-v4-pro.

Codex CLI

令牌选国产模型分组,编辑 ~/.codex/config.toml:

Create a key in the Chinese-models group and edit ~/.codex/config.toml:

model = "glm-5.3"
model_provider = "ainzy"

[model_providers.ainzy]
name = "Ainzy"
base_url = "https://ainzy.net/v1"
env_key = "AINZY_API_KEY"
wire_api = "responses"
model = "glm-5.3"
model_provider = "ainzy"

[model_providers.ainzy]
name = "Ainzy"
base_url = "https://ainzy.net/v1"
env_key = "AINZY_API_KEY"
wire_api = "responses"

然后设环境变量 AINZY_API_KEY=sk-你的key 再运行 codex。

Then set AINZY_API_KEY=sk-YOUR_KEY in your environment and run codex.

Cursor / Cline / Roo Code / 其他 OpenAI 兼容工具

Cursor / Cline / Roo Code / other OpenAI-compatible tools

  • 提供商选 OpenAI 或 OpenAI Compatible
  • Base URL 填 https://ainzy.net/v1,API Key 填你的 sk- key
  • 模型名手动填写分组里的模型,例如 deepseek-v4-flash、glm-5.3
  • Choose OpenAI or OpenAI Compatible as the provider
  • Base URL https://ainzy.net/v1, API key your sk- key
  • Type the model name from your group, e.g. deepseek-v4-flash, glm-5.3

Cursor 覆盖 OpenAI Base URL 后请点「Verify」,通过即表示 key 和地址都对。

In Cursor, click "Verify" after overriding the OpenAI base URL. A pass means both the key and the URL are correct.

Cherry Studio

桌面聊天客户端,不写代码、只想直接对话时用它最省事。最快的办法是一键导入:到令牌页,点你那条令牌的「聊天」,选 Cherry Studio,会唤起软件并自动填好地址和 key(需要本机已安装 Cherry Studio)。

A desktop chat client — the easiest option when you just want to talk to a model instead of writing code. Fastest path is the one-click import: on the keys page, open the "Chat" menu on your key and pick Cherry Studio. It launches the app with the base URL and key already filled in (Cherry Studio must already be installed).

手动配置同样可以:

Manual setup works just as well:

  • 设置 → 模型服务 → 添加提供商,类型选 OpenAI
  • API 地址填 https://ainzy.net,API 密钥填你的 sk- key
  • 在这个提供商下点「添加模型」,手动输入模型 ID,例如 glm-5.3、kimi-k3、deepseek-v4-pro
  • 回到聊天页,在顶部模型下拉里选中刚添加的模型
  • Settings → Model Providers → add a provider, type OpenAI
  • API host https://ainzy.net, API key your sk- key
  • Under that provider click "Add model" and type the model ID, e.g. glm-5.3, kimi-k3, deepseek-v4-pro
  • Back on the chat page, pick the new model from the dropdown at the top

模型下拉里空着,说明模型还没添加——Cherry Studio 不会自动列出本站模型,必须手填一次。若提示 404,把 API 地址改成 https://ainzy.net/v1/(末尾带斜杠)再试。

An empty model dropdown just means no model has been added yet — Cherry Studio does not list ours automatically, so add one by hand. If you get a 404, change the API host to https://ainzy.net/v1/ (with the trailing slash) and retry.

代码示例

Code samples

把 sk-你的key 换成你的令牌、模型名换成分组里的模型,直接运行。

Replace sk-YOUR_KEY with your key and the model with one from your group, then run.

curl https://ainzy.net/v1/chat/completions \
  -H "Authorization: Bearer sk-你的key" \
  -H "Content-Type: application/json" \
  -d '{
    "model": "deepseek-v4-flash",
    "messages": [{"role": "user", "content": "你好,介绍一下你自己"}],
    "stream": true
  }'
curl https://ainzy.net/v1/chat/completions \
  -H "Authorization: Bearer sk-YOUR_KEY" \
  -H "Content-Type: application/json" \
  -d '{
    "model": "deepseek-v4-flash",
    "messages": [{"role": "user", "content": "Hi, introduce yourself"}],
    "stream": true
  }'
from openai import OpenAI

client = OpenAI(base_url="https://ainzy.net/v1", api_key="sk-你的key")

stream = client.chat.completions.create(
    model="glm-5.3",
    messages=[{"role": "user", "content": "你好,介绍一下你自己"}],
    stream=True,
)
for chunk in stream:
    print(chunk.choices[0].delta.content or "", end="", flush=True)
from openai import OpenAI

client = OpenAI(base_url="https://ainzy.net/v1", api_key="sk-YOUR_KEY")

stream = client.chat.completions.create(
    model="glm-5.3",
    messages=[{"role": "user", "content": "Hi, introduce yourself"}],
    stream=True,
)
for chunk in stream:
    print(chunk.choices[0].delta.content or "", end="", flush=True)
import OpenAI from "openai";

const client = new OpenAI({ baseURL: "https://ainzy.net/v1", apiKey: "sk-你的key" });

const stream = await client.chat.completions.create({
  model: "kimi-k3",
  messages: [{ role: "user", content: "你好,介绍一下你自己" }],
  stream: true,
});
for await (const chunk of stream) {
  process.stdout.write(chunk.choices[0]?.delta?.content ?? "");
}
import OpenAI from "openai";

const client = new OpenAI({ baseURL: "https://ainzy.net/v1", apiKey: "sk-YOUR_KEY" });

const stream = await client.chat.completions.create({
  model: "kimi-k3",
  messages: [{ role: "user", content: "Hi, introduce yourself" }],
  stream: true,
});
for await (const chunk of stream) {
  process.stdout.write(chunk.choices[0]?.delta?.content ?? "");
}
import anthropic

client = anthropic.Anthropic(base_url="https://ainzy.net", api_key="sk-你的key")

with client.messages.stream(
    model="glm-5.3",
    max_tokens=1024,
    messages=[{"role": "user", "content": "你好,介绍一下你自己"}],
) as stream:
    for text in stream.text_stream:
        print(text, end="", flush=True)
import anthropic

client = anthropic.Anthropic(base_url="https://ainzy.net", api_key="sk-YOUR_KEY")

with client.messages.stream(
    model="glm-5.3",
    max_tokens=1024,
    messages=[{"role": "user", "content": "Hi, introduce yourself"}],
) as stream:
    for text in stream.text_stream:
        print(text, end="", flush=True)
curl https://ainzy.net/v1/responses \
  -H "Authorization: Bearer sk-你的key" \
  -H "Content-Type: application/json" \
  -d '{
    "model": "glm-5.3",
    "input": "用一句话解释什么是 API 网关"
  }'
curl https://ainzy.net/v1/responses \
  -H "Authorization: Bearer sk-YOUR_KEY" \
  -H "Content-Type: application/json" \
  -d '{
    "model": "glm-5.3",
    "input": "Explain what an API gateway is in one sentence"
  }'

所有示例都用国产模型分组的 key 即可,Responses 接口同样支持国产模型。

All samples work with a Chinese-model-group key; the Responses endpoint supports the Chinese models too.

计费说明

Billing

  • 按量计费,无月费。只按实际消耗的 token 扣费,输入、输出、缓存命中分别计价。
  • 美元计价。余额单位是美元,价格页显示每百万 token 的美元单价。支付宝按 ¥7 = $1 换算。
  • 分组决定折扣。同一模型在不同分组价格不同,令牌建好后分组即固定。
  • 缓存命中更便宜。国产模型分组对命中缓存的输入 token 按缓存价计费,重复上下文越多越省。
  • 失败请求不扣费。返回错误的请求不会计入账单。
  • 不限单账号并发。没有每账号并发上限,高并发直接跑即可。
  • 邀请返利 10%。在钱包页拿邀请链接,被邀请用户每笔充值的 10% 会在几分钟内进入你的邀请奖励,可划转到余额。
  • Pay as you go, no monthly fee. Only tokens actually consumed are charged. Input, output and cache-hit tokens are priced separately.
  • Priced in US dollars. Balance is in USD and the pricing page shows USD per million tokens. Alipay converts at ¥7 = $1.
  • The group sets the discount. The same model costs differently in different groups, and a key's group is fixed once created.
  • Cache hits cost less. In the Chinese-model groups, cached input tokens are billed at the cache rate. The more repeated context, the more you save.
  • Failed requests are free. Requests that return an error are not billed.
  • No per-account concurrency cap. Run as many parallel requests as you need.
  • 10% referral rebate. Get your invite link on the wallet page. 10% of every top-up by an invited user lands in your referral balance within minutes and can be moved to your main balance.

每一次调用的模型、token 数、扣费都记录在使用日志,账单逐条可查。

Every call's model, token counts and charge are recorded in the usage log, so the bill is auditable line by line.

错误对照

Errors

常见报错和处理方法。

Common errors and what to do about them.

状态Status返回内容Response原因与处理Cause and fix
401Invalid tokenkey 填错、没有 sk- 前缀、请求头不是 Bearer,或令牌已被禁用。到令牌页重新复制。Wrong key, missing sk- prefix, header not Bearer, or the key is disabled. Copy it again from the keys page.
503model_not_found
分组 … 下模型 … 无可用线路model … not available in group …
这把 key 的分组里没有这个模型。最常见的原因是工具还在用默认模型名(例如 Claude Code 的 claude-…)。检查模型名拼写,或换用对应分组的 key。用 /v1/models 可以看当前 key 到底能用哪些模型。The key’s group does not include this model. The usual cause is a tool still using its default model name (for example Claude Code’s claude-…). Check the spelling, or use a key from the right group. /v1/models lists exactly what this key can use.
403用户额度不足Insufficient quota
剩余额度: $…remaining: $…
余额不够本次请求的预估费用。到钱包充值后重试。Balance is below the estimated cost of this request. Top up in the wallet and retry.
429请求过于频繁Too many requests短时间请求量过大,退避几秒后重试即可。Too many requests in a short window. Back off a few seconds and retry.
5xxupstream_error模型侧临时故障,稍后重试。失败请求不计费。持续出现请联系我们。Temporary fault on the model side. Retry later. Failed requests are not billed. Contact us if it persists.

常见问题

FAQ

国产模型第一个字出来很慢?

部分国产模型在回答前会先完成思考,首字通常在 10 秒左右,之后输出很快。建议始终开启 stream: true,客户端超时设到 120 秒以上。

国产模型的 key 放进 CC Switch / Claude Code 用不了,报 model not found?

地址和 key 都对,问题出在模型名。Claude Code 默认要 claude- 开头的模型,而国产模型分组里只有 glm-5.3、kimi-k3、deepseek-v4-pro 这类名字,所以会报 model_not_found。在 CC Switch 里点一次「拉取模型」把模型名换掉,或者按上面 工具接入 里 Claude Code 那段,补上 ANTHROPIC_MODEL 和 ANTHROPIC_SMALL_FAST_MODEL 两个环境变量。不想折腾命令行的话,用 Cherry Studio 直接对话最省事。

一把 key 能不能同时用 Claude Code 和 Codex?

能。同一把国产模型分组的 key,Claude Code 走 Anthropic 协议、Codex 走 Responses 协议,都可以用。CC Switch 可以一键切换两套配置。

侧边栏「聊天」菜单点开是空白页?

「聊天」菜单里的 Cherry Studio、AionUI、DeepChat、AMA 问天、OpenCat 是桌面客户端的一键导入链接,需要先在本机装好对应客户端,点击后会直接唤起应用并自动填入本站地址和你的 key。没装客户端时浏览器只会开一个空白标签,属于正常现象。CC Switch 一项会打开本站的 CC Switch 指南;Lobe Chat 和 AI as Workspace 是网页版,直接可用。

调用日志里的费用和价格页对不上?

价格页显示的是每百万 token 的单价,日志里是本次请求的实际扣费,等于输入 token × 输入单价 + 输出 token × 输出单价 + 缓存命中 token × 缓存单价,再除以一百万。

支持函数调用、JSON 模式、视觉输入吗?

跟随模型本身的能力,网关原样透传请求参数。工具调用、JSON 输出、视觉输入等以各模型官方文档为准,参数写法与官方 API 一致,无需改动。

如何联系人工?

邮件 contact@ainzy.cn。写明账号邮箱、出错时间和使用日志里的那条记录,我们能更快定位。

The first token from a Chinese model is slow?

Some of these models think before they answer, so the first token usually arrives after about 10 seconds and the rest streams quickly. Always use stream: true and set the client timeout to 120 seconds or more.

My Chinese-models key fails in CC Switch / Claude Code with model not found

The base URL and key are fine — the model name is the problem. Claude Code asks for claude- models by default, while the Chinese-models groups only contain names such as glm-5.3, kimi-k3 and deepseek-v4-pro, hence model_not_found. Either click "fetch models" once in CC Switch to replace the model name, or follow the Claude Code block under Tools and set ANTHROPIC_MODEL and ANTHROPIC_SMALL_FAST_MODEL. If you would rather not touch a terminal, Cherry Studio is the simplest way to just chat.

Can one key serve both Claude Code and Codex?

Yes. One Chinese-model-group key works for both: Claude Code uses the Anthropic protocol and Codex uses the Responses protocol. CC Switch can flip between the two configs with one click.

An item in the sidebar "Chat" menu opens a blank page?

Cherry Studio, AionUI, DeepChat, AMA and OpenCat in that menu are one-click import links for desktop clients. Install the client first; clicking then launches the app with this site's address and your key filled in. Without the client installed the browser just opens a blank tab, which is expected. The CC Switch item opens our CC Switch guide; Lobe Chat and AI as Workspace are web apps and work directly.

The charge in the usage log does not match the pricing page?

The pricing page shows the rate per million tokens. The log shows the actual charge for one request: input tokens × input rate + output tokens × output rate + cached tokens × cache rate, divided by one million.

Are function calling, JSON mode and vision supported?

Whatever the model itself supports. The gateway passes request parameters through unchanged, so tool calling, JSON output and vision follow each model's official documentation, with the same parameter names as the official API.

How do I reach a human?

Email contact@ainzy.cn. Include your account email, the time of the request and the matching line from the usage log so we can find it quickly.

开始使用

Get started

注册账号,充值后建一把令牌就能调用

Register, add credit, create a key, and start calling

免费注册 →Create account →