{"data":[{"schema_version":"2.4","id":"zai-org/glm-5.3-flash","name":"GLM 5.3 Flash","input_modalities":[{"type":"text","supported_inputs":{"max_context_length":{"value":1048576,"unit":"token"}},"pricing":[{"type":"prompt","unit":"token","cost_usd":"0.000000110"}],"capacity":[{"type":"prompt","unit":"token","per":"minute","value":50000000}]},{"type":"image"},{"type":"video"}],"output_modalities":[{"type":"text","supported_parameters":{"temperature":{"type":"range","min":0,"max":2,"default":1},"top_p":{"type":"range","min":0,"max":1,"default":1},"max_completion_tokens":{"type":"integer","min":1,"unit":"token"},"reasoning":{"type":"boolean"},"tool_choice":{"type":"unknown"},"structured_outputs":{"type":"boolean"},"top_a":{"type":"range","min":0,"max":1,"default":0},"top_k":{"type":"integer","min":0,"default":0}},"max_length":{"value":128000,"unit":"token"},"streaming":true,"pricing":[{"type":"completion","unit":"token","cost_usd":"0.000000400"}],"capacity":[{"type":"completion","unit":"token","per":"minute","value":500000},{"type":"concurrency","unit":"request","value":100}]}],"capacity":[{"type":"request","unit":"request","per":"minute","value":100}],"created":1790748977,"quantization":"fp8","description":"GLM 5.3 Flash is well suited to coding agents, repository-scale debugging, tool-using workflows and visual tasks such as understanding screenshots, charts and documents. Use it when an agent needs to work across a large codebase or long research context, call tools, inspect visual inputs and continue through multi-step tasks without switching models. Z.ai positions the model particularly around coding, multimodal understanding and agentic work.","datacenters":[{"country_code":"US","region":"us"}],"compliance":{"zdr":true,"hipaa":false,"gdpr":true},"is_ready":true},{"schema_version":"2.4","id":"deepseek-ai/DeepSeek-V4.1-Flash","name":"DeepSeek V4.1 Flash","input_modalities":[{"type":"text","supported_inputs":{"max_context_length":{"value":1048576,"unit":"token"}},"pricing":[{"type":"prompt","unit":"token","cost_usd":"0.000000300"}]}],"output_modalities":[{"type":"text","supported_parameters":{"temperature":{"type":"range","min":0,"max":2,"default":1},"top_p":{"type":"range","min":0,"max":1,"default":1},"top_k":{"type":"integer","min":0,"default":0},"frequency_penalty":{"type":"range","min":-2,"max":2,"default":0},"max_tokens":{"type":"integer","min":1,"unit":"token","max":384000},"structured_outputs":{"type":"boolean"},"tools":{"type":"boolean","default":true},"reasoning":{"type":"boolean","default":true}},"max_length":{"value":384000,"unit":"token"},"streaming":true,"pricing":[{"type":"completion","unit":"token","cost_usd":"0.000001100"}]}],"hugging_face_id":"deepseek-ai/DeepSeek-V4.1-Flash","created":1790269073,"quantization":"fp8","description":"DeepSeek V4.1 Flash is built for fast coding, software-engineering agents and high-volume agentic workloads that still require strong reasoning. It is a good fit for terminal automation, code debugging, tool calling, research agents, document or screenshot analysis, and long-running workflows that repeatedly reuse context. For new DeepSeek Flash integrations, V4.1 is the newer multimodal generation and should normally replace the original V4 Flash.","datacenters":[{"country_code":"US","region":"us"}],"compliance":{"zdr":true,"hipaa":false,"gdpr":true},"openrouter":{"slug":"deepseek-ai/DeepSeek-V4.1-Flash"},"is_ready":true},{"schema_version":"2.4","id":"openai/gpt-oss-120b","name":"openai/gpt-oss-120b","input_modalities":[{"type":"text","supported_inputs":{"max_context_length":{"value":131072,"unit":"token"}},"pricing":[{"type":"prompt","unit":"token","cost_usd":"0.000000039"}],"capacity":[{"type":"prompt","unit":"token","per":"minute","value":50000000}]}],"output_modalities":[{"type":"text","supported_parameters":{"temperature":{"type":"range","min":0,"max":2,"default":1},"presence_penalty":{"type":"range","min":-2,"max":2,"default":0},"frequency_penalty":{"type":"range","min":-2,"max":2,"default":0},"seed":{"type":"integer"},"max_tokens":{"type":"integer","min":1,"unit":"token","max":131072},"stop":{"type":"array"},"repetition_penalty":{"type":"range","min":0,"max":2,"default":1},"top_k":{"type":"integer","min":0,"default":0},"tools":{"type":"boolean"},"response_format":{"type":"unknown"},"structured_outputs":{"type":"boolean"},"reasoning":{"type":"boolean","default":true}},"max_length":{"value":131072,"unit":"token"},"streaming":true,"pricing":[{"type":"completion","unit":"token","cost_usd":"0.000000158"}],"capacity":[{"type":"completion","unit":"token","per":"minute","value":500000},{"type":"concurrency","unit":"request","value":100}]}],"capacity":[{"type":"request","unit":"request","per":"minute","value":100}],"hugging_face_id":"openai/gpt-oss-120b","created":1769783160,"quantization":"bf16","description":"GPT-OSS-120B is OpenAI's higher-capability open-weight model for complex reasoning, coding agents, function calling and structured outputs. Use it for workloads that require multi-step problem solving, planning, tool use or more difficult coding tasks, especially when you want an open model that can be customized or deployed outside OpenAI's hosted API.","datacenters":[{"country_code":"US","region":"us"},{"country_code":"NO","region":"eu-no-stavanger"}],"compliance":{"zdr":true,"hipaa":false,"gdpr":true},"openrouter":{"slug":"openai/gpt-oss-120b"},"is_ready":true},{"schema_version":"2.4","id":"deepseek-ai/DeepSeek-V4-Flash","name":"DeepSeek V4 Flash","input_modalities":[{"type":"text","supported_inputs":{"max_context_length":{"value":1048576,"unit":"token"}},"pricing":[{"type":"prompt","unit":"token","cost_usd":"0.000000140"}],"capacity":[{"type":"prompt","unit":"token","per":"minute","value":50000000}]}],"output_modalities":[{"type":"text","supported_parameters":{"temperature":{"type":"range","min":0,"max":2,"default":1},"top_p":{"type":"range","min":0,"max":1,"default":1},"top_k":{"type":"integer","min":0,"default":0},"min_p":{"type":"range","min":0,"max":1,"default":0},"frequency_penalty":{"type":"range","min":-2,"max":2,"default":0},"presence_penalty":{"type":"range","min":-2,"max":2,"default":0},"max_tokens":{"type":"integer","min":1,"unit":"token","max":393216},"seed":{"type":"integer"},"stop":{"type":"array"},"tools":{"type":"boolean"},"response_format":{"type":"unknown"},"structured_outputs":{"type":"boolean"},"reasoning":{"type":"boolean"},"include_reasoning":{"type":"boolean","default":true}},"max_length":{"value":393216,"unit":"token"},"streaming":true,"pricing":[{"type":"completion","unit":"token","cost_usd":"0.000000300"}],"capacity":[{"type":"completion","unit":"token","per":"minute","value":500000},{"type":"concurrency","unit":"request","value":100}]}],"capacity":[{"type":"request","unit":"request","per":"minute","value":100}],"hugging_face_id":"deepseek-ai/DeepSeek-V4-Flash","created":1779196213,"quantization":"fp8","description":"DeepSeek V4 Flash is the speed-oriented DeepSeek V4 model for coding, reasoning and long-context agent workflows. It works well for code generation, repository analysis, automated research and tool-driven applications where throughput matters and the heavier Pro model is unnecessary. For new applications that need image understanding or DeepSeek's latest Flash capabilities, use V4.1 Flash instead.","datacenters":[{"country_code":"US","region":"us"}],"compliance":{"zdr":true,"hipaa":false,"gdpr":true},"openrouter":{"slug":"deepseek-ai/DeepSeek-V4-Flash"},"is_ready":true},{"schema_version":"2.4","id":"moonshotai/Kimi-K2.6","name":"Kimi K2.6","input_modalities":[{"type":"text","supported_inputs":{"max_context_length":{"value":262144,"unit":"token"}},"pricing":[{"type":"prompt","unit":"token","cost_usd":"0.000000600"}],"capacity":[{"type":"prompt","unit":"token","per":"minute","value":50000000}]}],"output_modalities":[{"type":"text","supported_parameters":{"temperature":{"type":"range","min":0,"max":2,"default":1},"top_p":{"type":"range","min":0,"max":1,"default":1},"top_k":{"type":"integer","min":0,"default":0},"presence_penalty":{"type":"range","min":-2,"max":2,"default":0},"repetition_penalty":{"type":"range","min":0,"max":2,"default":1},"seed":{"type":"integer"},"max_tokens":{"type":"integer","min":1,"unit":"token","max":262144},"stop":{"type":"array"},"tools":{"type":"boolean"},"response_format":{"type":"unknown"},"structured_outputs":{"type":"boolean"},"reasoning":{"type":"boolean"}},"max_length":{"value":262144,"unit":"token"},"streaming":true,"pricing":[{"type":"completion","unit":"token","cost_usd":"0.000003200"}],"capacity":[{"type":"completion","unit":"token","per":"minute","value":500000},{"type":"concurrency","unit":"request","value":100}]}],"capacity":[{"type":"request","unit":"request","per":"minute","value":100}],"hugging_face_id":"moonshotai/Kimi-K2.6","created":1778638064,"quantization":"fp8","description":"Kimi K2.6 is especially suited to coding-heavy agents that need to keep working across long sequences of actions. Use it for full-stack application development, frontend generation, debugging, autonomous research and workflows where multiple agents or subtasks need to be coordinated. It is particularly relevant to coding agents that must build, inspect, modify and test rather than simply generate a single code snippet.","datacenters":[{"country_code":"US","region":"us"}],"compliance":{"zdr":true,"hipaa":false,"gdpr":true},"openrouter":{"slug":"moonshotai/kimi-k2-thinking"},"is_ready":true},{"schema_version":"2.4","id":"zai-org/glm-5.2","name":"GLM 5.2","input_modalities":[{"type":"text","supported_inputs":{"max_context_length":{"value":1048576,"unit":"token"}},"pricing":[{"type":"prompt","unit":"token","cost_usd":"0.00000090"}],"capacity":[{"type":"prompt","unit":"token","per":"minute","value":50000000}]}],"output_modalities":[{"type":"text","supported_parameters":{"temperature":{"type":"range","min":0,"max":2,"default":1},"top_p":{"type":"range","min":0,"max":1,"default":1},"frequency_penalty":{"type":"range","min":-2,"max":2,"default":0},"max_tokens":{"type":"integer","min":1,"unit":"token","max":131072},"stop":{"type":"array"},"min_p":{"type":"range","min":0,"max":1,"default":0},"tools":{"type":"boolean"},"structured_outputs":{"type":"boolean"},"reasoning":{"type":"boolean"}},"max_length":{"value":131072,"unit":"token"},"streaming":true,"pricing":[{"type":"completion","unit":"token","cost_usd":"0.00000280"}],"capacity":[{"type":"completion","unit":"token","per":"minute","value":500000},{"type":"concurrency","unit":"request","value":100}]}],"capacity":[{"type":"request","unit":"request","per":"minute","value":100}],"hugging_face_id":"zai-org/glm-5.2","created":1782985175,"quantization":"fp8","description":"GLM 5.2 is designed for long-horizon software engineering and agent workflows where the model must stay coherent across a large amount of project context. It is a strong fit for multi-file refactoring, complex debugging, large-scale implementation, automated research, performance optimization and coding tasks that may continue for many agent steps.","datacenters":[{"country_code":"US","region":"us"}],"compliance":{"zdr":true,"hipaa":false,"gdpr":true},"openrouter":{"slug":"zai-org/glm-5.2"},"is_ready":true},{"schema_version":"2.4","id":"deepseek-ai/DeepSeek-V4-Pro","name":"DeepSeek V4 Pro","input_modalities":[{"type":"text","supported_inputs":{"max_context_length":{"value":1000000,"unit":"token"}},"pricing":[{"type":"prompt","unit":"token","cost_usd":"0.000001480"}],"capacity":[{"type":"prompt","unit":"token","per":"minute","value":50000000}]}],"output_modalities":[{"type":"text","supported_parameters":{"temperature":{"type":"range","min":0,"max":2,"default":1},"top_p":{"type":"range","min":0,"max":1,"default":1},"max_tokens":{"type":"integer","min":1,"unit":"token","max":80000},"seed":{"type":"integer"},"stop":{"type":"array"},"top_k":{"type":"integer","min":0,"default":0},"frequency_penalty":{"type":"range","min":-2,"max":2,"default":0},"presence_penalty":{"type":"range","min":-2,"max":2,"default":0},"repetition_penalty":{"type":"range","min":0,"max":2,"default":1},"min_p":{"type":"range","min":0,"max":1,"default":0},"structured_outputs":{"type":"boolean"},"tools":{"type":"boolean"},"response_format":{"type":"unknown"}},"max_length":{"value":80000,"unit":"token"},"streaming":true,"pricing":[{"type":"completion","unit":"token","cost_usd":"0.000003400"}],"capacity":[{"type":"completion","unit":"token","per":"minute","value":500000},{"type":"concurrency","unit":"request","value":100}]}],"capacity":[{"type":"request","unit":"request","per":"minute","value":100}],"hugging_face_id":"deepseek-ai/DeepSeek-V4-Pro","created":1777703520,"quantization":"fp8","description":"DeepSeek V4 Pro is the higher-compute V4 model for difficult reasoning, complex coding and knowledge-intensive agent workflows. Use it for architecture work, hard debugging, technical analysis and multi-step problems where you are willing to trade higher inference cost for deeper reasoning. For many newer workloads, also compare it with V4.1 Flash, which DeepSeek now positions as its newer high-efficiency option.","datacenters":[{"country_code":"US","region":"us"}],"compliance":{"zdr":true,"hipaa":false,"gdpr":true},"openrouter":{"slug":"deepseek-ai/DeepSeek-V4-Pro"},"is_ready":true},{"schema_version":"2.4","id":"google/gemma-4-31B-it","name":"Gemma 4 31B","input_modalities":[{"type":"text","supported_inputs":{"max_context_length":{"value":262144,"unit":"token"}},"pricing":[{"type":"prompt","unit":"token","cost_usd":"0.000000130"}],"capacity":[{"type":"prompt","unit":"token","per":"minute","value":50000000}]},{"type":"image"},{"type":"video"}],"output_modalities":[{"type":"text","supported_parameters":{"temperature":{"type":"range","min":0,"max":2,"default":1},"top_p":{"type":"range","min":0,"max":1,"default":1},"top_k":{"type":"integer","min":0,"default":0},"frequency_penalty":{"type":"range","min":-2,"max":2,"default":0},"presence_penalty":{"type":"range","min":-2,"max":2,"default":0},"repetition_penalty":{"type":"range","min":0,"max":2,"default":1},"max_tokens":{"type":"integer","min":1,"unit":"token","max":262144},"tools":{"type":"boolean"},"response_format":{"type":"unknown"},"structured_outputs":{"type":"boolean"}},"max_length":{"value":262144,"unit":"token"},"streaming":true,"pricing":[{"type":"completion","unit":"token","cost_usd":"0.000000370"}],"capacity":[{"type":"completion","unit":"token","per":"minute","value":500000},{"type":"concurrency","unit":"request","value":100}]}],"capacity":[{"type":"request","unit":"request","per":"minute","value":100}],"hugging_face_id":"google/gemma-4-31B-it","created":1779643266,"quantization":"fp4","description":"Gemma 4 31B is a general-purpose Google open model for reasoning, coding, multilingual assistants and applications that need to understand both text and images. It fits document analysis, visual question answering, knowledge assistants, code generation and applications serving users across many languages. It is particularly useful when you want capable multimodal reasoning without moving to one of the extremely large open MoE models.","datacenters":[{"country_code":"US","region":"us"}],"compliance":{"zdr":true,"hipaa":false,"gdpr":true},"openrouter":{"slug":"google/gemma-4-31b-it"},"is_ready":true},{"schema_version":"2.4","id":"openai/gpt-oss-20b","name":"openai/gpt-oss-20b","input_modalities":[{"type":"text","supported_inputs":{"max_context_length":{"value":131100,"unit":"token"}},"pricing":[{"type":"prompt","unit":"token","cost_usd":"0.000000030"}],"capacity":[{"type":"prompt","unit":"token","per":"minute","value":50000000}]}],"output_modalities":[{"type":"text","supported_parameters":{"temperature":{"type":"range","min":0,"max":2,"default":1},"top_p":{"type":"range","min":0,"max":1,"default":1},"reasoning":{"type":"boolean","default":true}},"max_length":{"value":131100,"unit":"token"},"streaming":true,"pricing":[{"type":"completion","unit":"token","cost_usd":"0.00000014"}],"capacity":[{"type":"completion","unit":"token","per":"minute","value":500000},{"type":"concurrency","unit":"request","value":100}]}],"capacity":[{"type":"request","unit":"request","per":"minute","value":100}],"hugging_face_id":"openai/gpt-oss-20b","created":1769783400,"quantization":"bf16","description":"GPT-OSS-20B is OpenAI's lower-latency open-weight reasoning model for frequent, well-defined developer and agent tasks. Use it for extraction, classification, routing, structured outputs, tool calls, lightweight coding and high-volume agent loops where responsiveness matters more than maximum reasoning capability. It is the natural GPT-OSS choice when 120B would be unnecessarily heavy for the task.","datacenters":[{"country_code":"US","region":"us"},{"country_code":"NO","region":"eu-no-stavanger"}],"compliance":{"zdr":true,"hipaa":false,"gdpr":true},"openrouter":{"slug":"openai/gpt-oss-20b"},"is_ready":true}]}