Guide
2,496 step-by-step guides to setting up AI models from OpenAI, Anthropic, Google, OpenRouter, and more in TypingMind with your own API key.
Search and filter guides
Showing 1,786-1,836 of 2,496 guides

GLM-4.5
Z.AI · Jul 28, 2025
Hybrid-reasoning GLM release that made the 4.5 line broadly useful

GLM-4.5-Flash
Z.AI · Jul 28, 2025
Efficient GLM model for fast reasoning, coding, and agent workflows

GLM-4.5-Air
Z.AI · Jul 28, 2025
Lighter GLM-4.5 variant for fast coding assistance and cheaper agents

GLM 4.5
Vercel AI Gateway · Jul 28, 2025
Hybrid-reasoning GLM release that made the 4.5 line broadly useful

GLM 4.5 Air
Vercel AI Gateway · Jul 28, 2025
Lighter GLM-4.5 variant for fast coding assistance and cheaper agents

Qwen Flash
Alibaba Cloud · Jul 28, 2025
Efficient Qwen model for fast chat, extraction, and high-volume workloads

Qwen3 Coder Flash
Alibaba Cloud · Jul 28, 2025
Qwen coding model for software agents, repository edits, and code reasoning

GLM-4.5
Novita · Jul 28, 2025
Flagship GLM model for hybrid reasoning, coding, and agentic engineering

Z.ai: GLM 4.5
OpenRouter · Jul 25, 2025
GLM-4.5 is our latest flagship foundation model, purpose-built for agent-based applications. It leverages a Mixture-of-Experts (MoE) architecture and supports a context length of up to 128k tokens. GLM-4.5 delivers significantly...

Z.ai: GLM 4.5 Air (free)
OpenRouter · Jul 25, 2025
GLM-4.5-Air is the lightweight variant of our latest flagship model family, also purpose-built for agent-centric applications. Like GLM-4.5, it adopts the Mixture-of-Experts (MoE) architecture but with a more compact parameter...

Z.ai: GLM 4.5 Air
OpenRouter · Jul 25, 2025
GLM-4.5-Air is the lightweight variant of our latest flagship model family, also purpose-built for agent-centric applications. Like GLM-4.5, it adopts the Mixture-of-Experts (MoE) architecture but with a more compact parameter...

Qwen: Qwen3 235B A22B Thinking 2507
OpenRouter · Jul 25, 2025
Qwen3-235B-A22B-Thinking-2507 is a high-performance, open-weight Mixture-of-Experts (MoE) language model optimized for complex reasoning tasks. It activates 22B of its 235B parameters per forward pass and natively supports up to 262,144...

Qwen3 235B A22B Thinking 2507
Synthetic · Jul 25, 2025
Qwen3 235B A22B Thinking 2507 from Synthetic - text input, 256,000 token context

Llama 3.3 Nemotron Super 49B v1.5
DeepInfra · Jul 25, 2025
Nemotron model for efficient reasoning, coding, and specialized AI agents
Qwen3-235B-A22B-Thinking-2507
Hugging Face · Jul 25, 2025
Qwen reasoning model for deliberate problem solving, math, and coding

Qwen3 235B A22B Instruct 2507 FP8
Together AI · Jul 25, 2025
Qwen instruction model for multilingual chat, reasoning, and tool use

Qwen3 235B A22b Thinking 2507
Novita · Jul 25, 2025
Qwen reasoning model for deliberate problem solving, math, and coding

Z.ai: GLM 4 32B
OpenRouter · Jul 24, 2025
GLM 4 32B is a cost-effective foundation language model. It can efficiently perform complex tasks and has significantly enhanced capabilities in tool use, online search, and code-related intelligent tasks. It...

Qwen: Qwen3 Coder 480B A35B (free)
OpenRouter · Jul 23, 2025
Qwen3-Coder-480B-A35B-Instruct is a Mixture-of-Experts (MoE) code generation model developed by the Qwen team. It is optimized for agentic coding tasks such as function calling, tool use, and long-context reasoning over...

Qwen: Qwen3 Coder 480B A35B
OpenRouter · Jul 23, 2025
Qwen3-Coder-480B-A35B-Instruct is a Mixture-of-Experts (MoE) code generation model developed by the Qwen team. It is optimized for agentic coding tasks such as function calling, tool use, and long-context reasoning over...

Qwen: Qwen3 Coder 480B A35B (exacto)
OpenRouter · Jul 23, 2025
Qwen3-Coder-480B-A35B-Instruct is a Mixture-of-Experts (MoE) code generation model developed by the Qwen team. It is optimized for agentic coding tasks such as function calling, tool use, and long-context reasoning over repositories. The model features 480 billion total parameters, with 35 billion active per forward pass (8 out of 160 experts). Pricing for the Alibaba endpoints varies by context length. Once a request is greater than 128k input tokens, the higher pricing is used.

Qwen 3 Coder 480B
Synthetic · Jul 23, 2025
Qwen 3 Coder 480B from Synthetic - text input, 256,000 token context

Qwen3-Coder 480B-A35B Instruct
AWS Bedrock · Jul 23, 2025
Open Qwen coding heavyweight for repository reasoning and agentic engineering

Qwen3 Coder 480B A35B Instruct Turbo
DeepInfra · Jul 23, 2025
Qwen coding model for software agents, repository edits, and code reasoning

Qwen3 Coder 480B A35B Instruct
DeepInfra · Jul 23, 2025
Qwen3 Coder 480B A35B Instruct from DeepInfra - text input, 262,144 token context
Qwen3-Coder-480B-A35B-Instruct
Hugging Face · Jul 23, 2025
Qwen coding model for software agents, repository edits, and code reasoning

Qwen3 Coder 480B A35B Instruct
Together AI · Jul 23, 2025
Legacy model retained for compatibility with older integrations

Qwen3 Coder Plus
Vercel AI Gateway · Jul 23, 2025
Hosted Qwen coder for software agents, repo edits, and long-context code

Qwen3 Coder Plus
Alibaba Cloud · Jul 23, 2025
Hosted Qwen coder for software agents, repo edits, and long-context code

Qwen3 Coder 480B A35B Instruct
Novita · Jul 23, 2025
Qwen coding model for software agents, repository edits, and code reasoning

ByteDance: UI-TARS 7B
OpenRouter · Jul 22, 2025
UI-TARS-1.5 is a multimodal vision-language agent optimized for GUI-based environments, including desktop interfaces, web browsers, mobile systems, and games. Built by ByteDance, it builds upon the UI-TARS framework with reinforcement...

Google: Gemini 2.5 Flash Lite
OpenRouter · Jul 22, 2025
Gemini 2.5 Flash-Lite is a lightweight reasoning model in the Gemini 2.5 family, optimized for ultra-low latency and cost efficiency. It offers improved throughput, faster token generation, and better performance...

Google: Gemini 2.5 Flash Lite (batch)
OpenRouter · Jul 22, 2025
Gemini 2.5 Flash-Lite is a lightweight reasoning model in the Gemini 2.5 family, optimized for ultra-low latency and cost efficiency. It offers improved throughput, faster token generation, and better performance...

Qwen3 Coder 480B A35B Instruct
Fireworks AI · Jul 22, 2025
Qwen3 Coder 480B A35B Instruct from Fireworks AI - text input, 256,000 token context

Qwen3 Coder 480B A35B Instruct
Vercel AI Gateway · Jul 22, 2025
Qwen coding model for software agents, repository edits, and code reasoning

Qwen3 235B A22B Instruct 2507
Novita · Jul 22, 2025
Qwen instruction model for multilingual chat, reasoning, and tool use

Qwen: Qwen3 235B A22B Instruct 2507
OpenRouter · Jul 21, 2025
Qwen3-235B-A22B-Instruct-2507 is a multilingual, instruction-tuned mixture-of-experts language model based on the Qwen3-235B architecture, with 22B active parameters per forward pass. It is optimized for general-purpose text generation, including instruction following,...

Qwen3 235B-A22B Instruct 2507
AWS Bedrock · Jul 21, 2025
Updated large open Qwen3 MoE instruct model for multilingual chat, coding, and tool use

Qwen3 235B-A22B Instruct 2507
DeepInfra · Jul 21, 2025
Updated large open Qwen3 MoE instruct model for multilingual chat, coding, and tool use
Qwen3 235B-A22B Instruct 2507
Hugging Face · Jul 21, 2025
Qwen instruction model for multilingual chat, reasoning, and tool use

Voxtral Small (latest)
Mistral · Jul 15, 2025
Instruct model with native audio input for speech understanding and tool use

Voxtral Small 24B 2507
AWS Bedrock · Jul 15, 2025
Open audio-language model for speech transcription, audio understanding, and voice-driven tool use

Voxtral Mini 3B 2507
AWS Bedrock · Jul 15, 2025
Open audio-language model for speech transcription, audio understanding, and voice-driven tool use

Kimi K2 Instruct
Groq · Jul 14, 2025
Kimi K2 Instruct from Groq - text input, 131,072 token context

Kimi K2 0711
Moonshot AI · Jul 14, 2025
Kimi model for long-context chat, coding, and agentic reasoning
Kimi-K2-Instruct
Hugging Face · Jul 14, 2025
Kimi model for long-context chat, coding, and agentic reasoning

Switchpoint Router
OpenRouter · Jul 11, 2025
Switchpoint AI's router instantly analyzes your request and directs it to the optimal AI from an ever-evolving library. As the world of LLMs advances, our router gets smarter, ensuring you...

MoonshotAI: Kimi K2 0711 (free)
OpenRouter · Jul 11, 2025
Kimi K2 Instruct is a large-scale Mixture-of-Experts (MoE) language model developed by Moonshot AI, featuring 1 trillion total parameters with 32 billion active per forward pass. It is optimized for agentic capabilities, including advanced tool use, reasoning, and code synthesis. Kimi K2 excels across a broad range of benchmarks, particularly in coding (LiveCodeBench, SWE-bench), reasoning (ZebraLogic, GPQA), and tool-use (Tau2, AceBench) tasks. It supports long-context inference up to 128K tokens and is designed with a novel training stack that includes the MuonClip optimizer for stable large-scale MoE training.

MoonshotAI: Kimi K2 0711
OpenRouter · Jul 11, 2025
Kimi K2 Instruct is a large-scale Mixture-of-Experts (MoE) language model developed by Moonshot AI, featuring 1 trillion total parameters with 32 billion active per forward pass. It is optimized for...

THUDM: GLM 4.1V 9B Thinking
OpenRouter · Jul 11, 2025
GLM-4.1V-9B-Thinking is a 9B parameter vision-language model developed by THUDM, based on the GLM-4-9B foundation. It introduces a reasoning-centric "thinking paradigm" enhanced with reinforcement learning to improve multimodal reasoning, long-context understanding (up to 64K tokens), and complex problem solving. It achieves state-of-the-art performance among models in its class, outperforming even larger models like Qwen-2.5-VL-72B on a majority of benchmark tasks.

Kimi K2 Instruct
Fireworks AI · Jul 11, 2025
Kimi K2 Instruct from Fireworks AI - text input, 128,000 token context

