Model Marketplace
48 models available · Tiered specs with platform-price discounts
Showing 48 of 48 models
Claude Fable 5
Claude Fable 5 is Anthropic's most capable widely released model, built for demanding reasoning and long-horizon agentic work. Adaptive thinking is always on; 1M context, up to 128k output.
Claude Fable 5.1
Claude Fable 5.1 is an incremental update to Anthropic's most capable widely released model, built for demanding reasoning and long-horizon agentic work. Adaptive thinking is always on; 1M context, up to 128k output.
Claude Haiku 4.5
Claude Haiku 4.5 is Anthropic's fastest and most efficient model, delivering near-frontier intelligence at a fraction of the cost and latency of larger models.
Claude Opus 4.5
Claude Opus 4.5 is Anthropic's frontier reasoning model optimized for complex software engineering, agentic workflows, and long-horizon computer use.
Claude Opus 4.6
Opus 4.6 is Anthropic's strongest model for coding and long-running professional tasks, with frontier performance on agentic workflows and computer use.
Claude Opus 4.7
Opus 4.7 is the next generation of Anthropic's Opus family, built for long-running, asynchronous agentic workflows and sustained multi-step software engineering.
Claude Opus 4.8
Claude Opus 4.8 is Anthropic's most capable generally available model in the Opus family. It supports text, image, and file inputs with text output, with reasoning support and a 1M-token context window. Built for complex reasoning, long-horizon agentic coding, and high-autonomy professional work.
Claude Opus 5
Anthropic 旗舰模型,擅长复杂推理、编程与长链路 agentic 任务
Claude Sonnet 4.6
Sonnet 4.6 is Anthropic's most capable Sonnet-class model yet, with frontier performance across coding, agents, and professional knowledge work.
Claude Sonnet 5
The best combination of speed and intelligence
DeepSeek V4 Flash
DeepSeek-V4-Flash 是 DeepSeek 高效 MoE 模型,总参数 284B、激活 13B,支持 100 万 token 上下文,默认思考模式,兼顾推理速度与成本;支持 JSON 输出、Tool Calls、FIM 补全(非思考模式)。
DeepSeek V4 Pro
DeepSeek-V4-Pro 是 DeepSeek 旗舰 MoE 模型,支持 100 万 token 上下文、最大输出 384K,默认思考模式,面向复杂推理、编程与 Agent 任务;支持 JSON 输出、Tool Calls、上下文缓存。
DeepSeek V4.1 Flash
DeepSeek V4.1 Flash,552B MoE架构,支持思考模式与图像理解,1M上下文,384K最大输出。MIT开源。
Doubao Seed 2.1 Pro
Doubao Seed 2.1 Pro 是字节跳动豆包新一代旗舰深度思考模型,256K 上下文,面向复杂代码交付、长链路 Agent 执行与多模态视觉理解;在多项编程与科学计算基准上接近或超越 GPT-5.5。
Doubao Seed 2.1 Turbo
Doubao Seed 2.1 Turbo 是 Seed 2.1 系列高性价比版本,功能与 Pro 对齐、价格约为一半,适合高并发企业流量与低延迟场景。
Doubao Seedance 2.0
字节跳动 Seedance 2.0 旗舰版视频生成模型,采用统一多模态音视频联合生成架构,支持文字、图片、音频、视频四种模态输入与全模态参考,最高 4K 10bit 输出,15 秒多镜头音视频,可用率、物理准确度与一致性达工业级水准。
Doubao Seedance 2.0 Fast
字节跳动 Seedance 2.0 极速版。音视图文参考,更快,继承 2.0 核心优势。最高 720p,4-15 秒输出,适合大批量出片与快速试样。
Doubao Seedance 2.5
Doubao Seedance 2.5 是字节跳动推出的新一代视频生成模型,支持文本、图片、音频和视频等多模态参考输入,在长叙事、全模态参考和视频编辑能力上进一步提升,适用于高质量视频生成、多模态参考创作与视频编辑等场景。
Doubao-Seed-Evolving
Seed 最新 Coding & Agent 模型,持续升级。面向 Agent 与 Coding 场景打造,具备复杂任务编排、长程规划、代码生成与工具调用等核心能力。
ERNIE 5.0
百度 ERNIE 5.0 大语言模型,含 Thinking Preview 版本。
ERNIE 5.1
百度 ERNIE 5.1 大语言模型,支持长文本理解与生成。
GLM-5
GLM-5 是智谱最新模型,专为编程与智能体场景打造,擅长复杂系统工程与长程任务,支持 200K 上下文。
GLM-5.1
GLM-5.1 在编码能力方面实现了重大飞跃,尤其在处理长周期任务方面表现突出。与以往以微小交互为基础的模型不同,GLM-5.1 可以独立且持续地完成单一任务超过8小时,在整个过程中自主规划、执行并不断优化自身,最终交付完整且达到工程级标准的结果。 GLM-5.1 delivers a major leap in coding capability, with particularly significant gains in handling long-horizon tasks. Unlike previous models built around minute-level interactions, GLM-5.1 can work independently and continuously on a single task for more than 8 hours, autonomously planning, executing, and improving itself throughout the process, ultimately delivering complete, engineering-grade results.
GLM-5.2
GLM 5.2 是 Z.ai 开发的一款大规模推理模型,支持文本输入和输出,上下文窗口容量为 100 万个标记,适用于长周期的智能体工作流、项目级软件工程以及复杂的多步骤自动化任务。 推理能力高且支持xhigh;xhigh对应最高水平的推理。在长期任务中的编码和工具使用方面表现尤为出色,能够在整个开发流程中保持工程背景,并从需求到跨平台部署,始终一致地遵循标准,完成单一任务。 GLM 5.2 is a large-scale reasoning model from Z.ai. It supports text input and output with a 1M-token context window, and is suited for long-horizon agent workflows, project-level software engineering, and complex multi-step automation. Reasoning efforts high and xhigh are supported; xhigh maps to max reasoning. It is particularly strong at coding and tool use across long-running tasks, able to maintain engineering context and follow standards consistently through a full development workflow, from requirements to multi-platform deployment, in a single task.
GLM-5.3
GLM-5.3 是 Z.ai 最新旗舰基础模型,与 GLM-5.2 共用同一基座,全部提升来自后训练。面向复杂软件工程与智能体任务:在 Z.ai Code Bench 上较 GLM-5.2 提升 50%,Terminal-Bench 3.0 由 4.6 升至 28.3,Agents' Last Exam 由 23.8 升至 28.5,取得开源 SOTA;网络安全方向 CyberGym 得分 84.5%,ExploitBench 由 24.4% 升至 54.4%。支持 1M 上下文与 128K 最大输出,具备思考模式、函数调用、上下文缓存、结构化输出与 MCP 集成。
GLM-5.3-Flash
GLM-5.3的轻量高速版本,200K上下文,支持图片/视频/文件/文本多模态输入。
GPT-5.4
GPT-5.4 is OpenAI's latest frontier model, unifying the Codex and GPT lines into a single system with strong agentic and long-context capabilities.
GPT-5.4 Mini
GPT-5.4 mini brings the core capabilities of GPT-5.4 to a faster, more efficient model optimized for lower latency and cost.
GPT-5.5
GPT-5.5 is OpenAI's frontier model for complex reasoning and coding. Start here for the hardest professional workloads; supports text and image input with a 1M context window.
GPT-5.6 Luna
GPT-5.6 Luna 是 GPT-5.6 系列的轻量档,支持百万 Token 上下文,按提示长度分两档计价,单价为系列内最低,适合高并发、成本敏感的批量调用。
GPT-5.6 Sol
GPT-5.6 Sol 是 GPT-5.6 系列的高配档,支持百万 Token 上下文,按提示长度分两档计价,面向复杂推理与长链路任务。
GPT-5.6 Terra
GPT-5.6 Terra 是 GPT-5.6 系列的中间档,支持百万 Token 上下文,按提示长度分两档计价,在能力与单价之间取平衡,适合大多数生产工作流。
Gemini 2.5 Flash Lite
Gemini 2.5 Flash-Lite is the fastest and most budget-friendly multimodal model in the 2.5 family.
Gemini 3 Flash Preview
Gemini 3 Flash Preview is a high speed, high value thinking model designed for agentic workflows and multimodal tasks.
Gemini 3.1 Flash Lite
Gemini 3.1 Flash Lite is Google's GA high-efficiency multimodal model optimized for low-latency, high-volume tasks.
Gemini 3.1 Pro Preview
Gemini 3.1 Pro Preview delivers advanced intelligence, complex problem-solving, and powerful agentic and vibe coding capabilities.
Kimi K2.6
kimi/kimi-k2.6 是 Kimi 最新最智能的模型,月之暗面直供,在通用 Agent、代码、视觉理解等综合能力得到全面提升,同时支持文本、图片与视频输入、思考与非思考模式、对话与 Agent 任务。
Kimi K2.7 Code
月之暗面直供的 Kimi K2.7 Code 模型,以编码为中心的智能体模型,专为长程软件工程任务优化,仅支持思考模式。擅长跨多文件重构、功能实现、长会话调试等复杂工作流。
Kimi K3
Kimi K3 旗舰模型,2.8万亿参数,KDA混合线性注意力,1M上下文,原生视觉理解
MiniMax M2.5
顶尖性能与极致性价比,轻松驾驭复杂任务,输出速度约 60 TPS
MiniMax M2.7
开启模型的自我迭代,输出速度约 60 TPS
MiniMax M3
最新 M 系列语言模型,适用于 Agent 推理、工具调用、代码和长上下文任务(1M 上下文)
Qwen: Qwen3.6 Flash
Qwen3.6 原生视觉语言 Flash 系列模型,在整体性能上较 Qwen3.5-Flash 显著提升。重点增强了智能体编程能力(在多项代码智能体基准上大幅超越前代)、数学推理和代码推理能力;在视觉能力方面,空间智能显著增强,其中物体定位和目标检测表现尤为突出。
Qwen: Qwen3.6 Plus
Qwen3.6 Plus 原生视觉语言系列模型,在 Agentic coding、STEM 推理与视觉理解方面全面提升,适合生产工作流与高吞吐场景。
Qwen: Qwen3.7 Max
Qwen3.7-Max 是通义千问最新一代旗舰大模型,面向 Agent 场景全新设计,在长链路推理、跨文件代码理解与复杂工程任务执行等维度均有显著提升;支持百万 Token 上下文,默认开启思考模式,在编程、办公与生产力、长周期自主执行方面均能出色胜任各项任务。
Qwen: Qwen3.8 Max
Qwen3.8-Max 是通义千问最新一代旗舰多模态模型,支持文本、图像与视频输入、百万 Token 上下文与 128K 长输出,面向复杂 Agent 编排与长周期自主执行任务。
Qwen:Qwen3.7 Plus
在强大文本能力的基础上全面升级了视觉-语言能力,同时保持了在编码、工具使用和生产力工作流方面的完整智能体能力。其核心特色为多模态交互混合智能体能力,能够感知真实世界场景、读取屏幕并操作 GUI、基于视觉参考生成代码、端到端导航移动应用
Qwen:Qwen3.8 flash
通义千问3.8系列轻量高速模型,适合低延迟、高并发场景。