Image Generations
Date
Price

doubao-seed3d-2-0-260328

ByteDance’s next-generation, disruptive, high-precision, simulation-grade 3D asset generation large model
API
Image Processing
Pricing:
$11/1M Tokens

doubao-seed-2-1-turbo-260628

ByteDance’s new-generation, low-cost, low-latency, deep-thinking large model designed for large-scale production scenarios
Model
LLM
Model capability: imageModel capability: videoModel capability: thinkingModel capability: function_call
Input:
$0.43/1M tokens
Output:
$2.14/1M tokens

doubao-seed-2-1-pro-260628

ByteDance’s next-generation flagship model, designed for the era of Coding and Agents, featuring deep thinking capabilities.
Model
LLM
Model capability: imageModel capability: videoModel capability: thinkingModel capability: function_call
Input:
$0.86/1M tokens
Output:
$4.28/1M tokens

glm-5.2

Zhipu AI’s next-generation, self-developed, high-end flagship large model specifically designed for multimodal agents.
Model
LLM
Model capability: thinkingModel capability: function_call
Input:
$1.4/1M tokens
Output:
$4.4/1M tokens

kimi-k2.7-code

Kimi’s next-generation flagship large model for adaptive code engineering based on deep reinforcement learning
Model
LLM
Model capability: imageModel capability: videoModel capability: thinkingModel capability: function_call
Input:
$0.95/1M tokens
Output:
$4/1M tokens

qwen3.7-plus

Tongyi Qianwen 3.7: The Most Cost-Effective Multimodal Intelligent Agent Foundation in the Family
Model
LLM
Model capability: imageModel capability: videoModel capability: thinkingModel capability: function_call
Input:
$0.228/1M tokensstarting from
Output:
$0.92/1M tokensstarting from

MiniMax-M3

MiniMax’s new-generation trillion-parameter MoE multimodal flagship large model.
Model
LLM
Model capability: audioModel capability: imageModel capability: videoModel capability: thinkingModel capability: function_call
Input:
$0.6/1M tokensstarting from
Output:
$2.4/1M tokensstarting from

step-3.7-flash

[30-Day Limited-Time Free] StepStar’s Flagship Language Reasoning Model
Model
LLM
Model capability: imageModel capability: videoModel capability: thinkingModel capability: function_call
Pricing:
Limited Time Free

qwen3.7-max

Alibaba’s Qwen 3.7, a closed-source flagship large model designed specifically for the era of intelligent agents.
Model
LLM
Model capability: imageModel capability: videoModel capability: thinkingModel capability: function_call
Input:
$1.8/1M tokens
Output:
$5.3/1M tokens

happyhorse-1.0-r2v

Alibaba Group’s next-generation cutting-edge AI video generation model
API
Video Generation
Pricing:
$0.156/second

starting from

happyhorse-1.0-i2v

Alibaba Group’s next-generation cutting-edge AI video generation model
API
Video Generation
Pricing:
$0.156/second

starting from

happyhorse-1.0-t2v

Alibaba Group’s next-generation cutting-edge AI video generation model
API
Video Generation
Pricing:
$0.156/sec

starting from

deepseek-v4-pro

The latest flagship AI model released by the DeepSeek series represents the current highest standard in both scale and performance among open-source models.
Model
LLM
Model capability: thinkingModel capability: function_call
Input:
$0.43/1M tokens
Output:
$0.86/1M tokens

deepseek-v4-flash

DeepSeek’s newly released language model, designed for high-performance production scenarios, is specifically optimized for ultimate inference efficiency and response speed.
Model
LLM
Model capability: thinkingModel capability: function_call
Input:
$0.14/1M tokens
Output:
$0.28/1M tokens

kimi-k2.6

Kimi K2.6 is Kimi's newest and most intelligent model.
Model
LLM
Model capability: imageModel capability: videoModel capability: thinkingModel capability: function_call
Input:
$0.95/1M tokens
Output:
$4/1M tokens

qwen3.6-flash

The Qwen3.6 native vision-language series Flash model delivers significantly improved performance compared to the 3.5-Flash model.
Model
LLM
Model capability: imageModel capability: videoModel capability: thinkingModel capability: function_call
Input:
$0.19/1M tokensstarting from
Output:
$1.13/1M tokensstarting from

qwen3.6-35b-a3b

The Qwen3.6 series 35B-A3B native vision-language model, designed based on a hybrid architecture, integrates linear attention mechanisms with sparse mixture-of-experts models.
Model
LLM
Model capability: imageModel capability: videoModel capability: thinkingModel capability: function_call
Input:
$0.28/1M tokens
Output:
$1.7/1M tokens

wan2.7-videoedit

The video editing features of the 2.7 series consistently preserve detailed information such as the image subject, style, and text.
API
Video Generation
Pricing:
$0.1/second

starting from