Skip to main content

Model Profiles

Schatzi AI hosts all models exclusively on Swiss infrastructure, ensuring your data never leaves Switzerland. This page provides detailed technical specifications, capabilities, and pricing for each available model to help you select the optimal solution for your business requirements.

For guidance on selecting the right model for your use case, see /ai-models/choosing-model. To understand how token pricing works, visit /subscription-billing/understanding-tokens.

Apertus Swiss LLM - Large

Description

Designed for organizations requiring maximum transparency and regulatory compliance, this 70B parameter model serves as a reliable foundation for multilingual services, government applications, and research initiatives. With fully documented training methodologies and strict adherence to AI Act requirements, it prioritizes data privacy and intellectual property protection while delivering frontier-level performance.

Specifications

  • Context window: 65,536 tokens
  • Max output tokens: Not specified
  • Vision support: No
  • Reasoning mode: No
  • Function calling: No
  • Streaming: Yes
  • Availability: Chat UI & API

Ideal Use Cases

  • Government service automation and citizen support
  • Regulatory compliance documentation and reporting
  • Academic research and R&D documentation analysis
  • Multilingual public sector chatbots
  • Legal document review with transparency requirements
  • Cross-border administrative processes

Strengths

  • Complete training transparency and methodology documentation
  • Full AI Act compliance for regulated industries
  • Robust multilingual capabilities for European languages
  • 70B parameter scale delivering high performance
  • Guaranteed Swiss data sovereignty

Limitations

  • No vision or image analysis capabilities
  • No function calling support for tool integration

Pricing

  • Input: ... per million tokens
  • Output: ... per million tokens

When to Use

Choose this model when regulatory compliance, training transparency, and data sovereignty are non-negotiable requirements. It is specifically engineered for government agencies, research institutions, and organizations handling sensitive information that must remain within Swiss jurisdiction.

When to Choose a Different Model

For applications requiring vision capabilities or image analysis, use Chat & Document Analysis & Reasoning - Large or Chat, Document Analysis & Agent tasks - Xtra Large. If your workflow requires function calling or autonomous agent capabilities, select Reasoning & Agent tasks - Large.


Apertus Swiss LLM - Large (v1.5)

Description

The v1.5 iteration of the Apertus Swiss LLM provides a refined approach to multilingual services and R&D. It maintains the core commitment to transparency and AI Act compliance, offering a reliable and adaptable 70B parameter model for highly regulated environments.

Specifications

  • Context window: 65,536 tokens
  • Max output tokens: Not specified
  • Vision support: No
  • Reasoning mode: No
  • Function calling: No
  • Streaming: Yes
  • Availability: API only

Ideal Use Cases

  • High-precision multilingual translation
  • Government agency internal knowledge bases
  • R&D data synthesis
  • Compliant content generation for public sectors
  • Privacy-first corporate communications
  • Technical documentation in multiple languages

Strengths

  • Documented data and methods for full transparency
  • Strict adherence to privacy and intellectual property laws
  • High-performance 70B parameter architecture
  • Optimized for Swiss regulatory landscapes
  • Reliable multilingual output

Limitations

  • No vision or multimodal support
  • No function calling capabilities

Pricing

  • Input: ... per million tokens
  • Output: ... per million tokens

When to Use

Use this version when you require the specific refinements of the v1.5 architecture while maintaining the strict transparency and compliance standards of the Apertus series.

When to Choose a Different Model

If you need a model available in the Chat UI, use the standard Apertus Swiss LLM - Large. For tasks requiring vision, consider Chat, Vision, Document Analysis & Reasoning - Medium.


Document Analysis - Small

Description

A versatile multimodal model optimized for handling both text and image inputs to generate precise text outputs. It is designed for multilingual dialogue and document understanding, providing an efficient way to process visual information without the overhead of larger models.

Specifications

  • Context window: 32,000 tokens
  • Max output tokens: Not specified
  • Vision support: Yes
  • Reasoning mode: No
  • Function calling: No
  • Streaming: Yes
  • Availability: API only

Ideal Use Cases

  • Multilingual visual chat interfaces
  • Basic OCR and document summarization
  • Image captioning for accessibility
  • Visual content tagging
  • Simple document-based Q&A
  • Multilingual image-to-text translation

Strengths

  • Strong multimodal input handling
  • Optimized for multilingual dialogue
  • Efficient token usage for visual tasks
  • Fast response times
  • Reliable text generation from visual cues

Limitations

  • No function calling support
  • Limited context window for very long documents
  • No advanced reasoning capabilities

Pricing

  • Input: ... per million tokens
  • Output: ... per million tokens

When to Use

Choose this model for straightforward multimodal tasks where you need to describe images or answer questions about documents in multiple languages, but do not need to integrate with external APIs.

When to Choose a Different Model

If you require function calling to send extracted data to another system, use Chat, Multi-lingual, Coding & function calling - Small. For high-precision OCR of complex tables, use Document Analysis & OCR - Small (DeepSeek OCR).


Document Analysis - Xtra Small

Description

The most lightweight vision-language model in the series, optimized for maximum efficiency and low latency. It is designed for compact applications that require basic visual understanding and text generation with minimal cost.

Specifications

  • Context window: 16,384 tokens
  • Max output tokens: Not specified
  • Vision support: Yes
  • Reasoning mode: No
  • Function calling: No
  • Streaming: Yes
  • Availability: API only

Ideal Use Cases

  • High-speed image classification
  • Basic visual tagging for large datasets
  • Simple OCR for short snippets of text
  • Low-latency visual chat triggers
  • Mobile-optimized visual analysis
  • Basic document verification

Strengths

  • Lowest cost for vision-enabled tasks
  • Extremely fast inference and streaming
  • Low resource footprint
  • Efficient for simple, repetitive visual tasks
  • High throughput for batch processing

Limitations

  • Very limited context window
  • Lowest reasoning capability in the vision suite
  • Not suitable for complex document analysis

Pricing

  • Input: ... per million tokens
  • Output: ... per million tokens

When to Use

Deploy this model for high-volume, low-complexity visual tasks where speed and cost are the primary drivers and the input data is relatively small.

When to Choose a Different Model

For any task requiring complex analysis of a full page or multi-page document, move up to Chat, Vision, Document Analysis & Reasoning - Medium.


Fast Reasoning & Instruction Following - Small

Description

A specialized model optimized for high-speed reasoning and strict adherence to complex instructions. It is designed for developers who need a model that can follow precise formatting rules and logical constraints without the latency of larger reasoning models.

Specifications

  • Context window: 32,768 tokens
  • Max output tokens: Not specified
  • Vision support: No
  • Reasoning mode: No
  • Function calling: Yes
  • Streaming: Yes
  • Availability: API only

Ideal Use Cases

  • Structured data extraction (JSON/XML)
  • Strict template-based content generation
  • Fast logical validation of text
  • Instruction-heavy automation tasks
  • API response formatting
  • Rapid data analysis and categorization

Strengths

  • Exceptional instruction-following accuracy
  • Fast reasoning for simple to medium tasks
  • Full function calling support
  • Reliable structured output
  • High efficiency for developer workflows

Limitations

  • No vision capabilities
  • Not designed for long-form creative writing
  • Limited deep "thinking" for highly abstract problems

Pricing

  • Input: ... per million tokens
  • Output: ... per million tokens

When to Use

Use this model when your primary requirement is that the AI follows a specific set of rules or a strict format perfectly and quickly. It is ideal for the "glue" in an automation pipeline.

When to Choose a Different Model

For tasks requiring deep, multi-step logical deduction, use Reasoning & Problem Solving - Medium. For multimodal tasks, use Chat, Vision, Document Analysis & Reasoning - Medium.


Reasoning & Problem Solving - Small

Description

An entry-level reasoning model optimized for "thinking" and logical problem solving. It utilizes a reasoning process to work through problems step-by-step, providing more reliable answers for logical queries than standard chat models.

Specifications

  • Context window: 32,768 tokens
  • Max output tokens: Not specified
  • Vision support: No
  • Reasoning mode: Yes
  • Function calling: Yes
  • Streaming: Yes
  • Availability: API only

Ideal Use Cases

  • Basic mathematical problem solving
  • Logical puzzle resolution
  • Simple code debugging
  • Step-by-step instructional generation
  • Basic analytical reasoning
  • Logical consistency checking

Strengths

  • Native reasoning capabilities
  • Cost-effective entry point for "thinking" models
  • Supports function calling for tool integration
  • Higher accuracy on logical tasks than standard LLMs
  • Efficient streaming of reasoning steps

Limitations

  • Limited capacity for extremely complex architectural problems
  • No vision support
  • Smaller context window than document models

Pricing

  • Input: ... per million tokens
  • Output: ... per million tokens

When to Use

Choose this model for tasks that require a basic level of logical deduction or step-by-step thinking where a standard chat model might hallucinate, but where the highest level of reasoning is not required.

When to Choose a Different Model

For highly complex scientific or mathematical problems, upgrade to Reasoning & Problem Solving - Xtra Large. For agentic tasks, use Reasoning & Agent tasks - Large.


Llama 4 Maverick multi modal - Small

Description

A cutting-edge multimodal model optimized for seamless experiences across text and visual inputs. It is designed to understand the relationship between images and text, providing a fluid interface for multimodal applications.

Specifications

  • Context window: 32,768 tokens
  • Max output tokens: Not specified
  • Vision support: Yes
  • Reasoning mode: No
  • Function calling: Yes
  • Streaming: Yes
  • Availability: API only

Ideal Use Cases

  • Multimodal AI assistants
  • Visual content analysis for social media
  • Image-based product support
  • Interactive visual storytelling
  • Multimodal data entry automation
  • Visual Q&A for e-commerce

Strengths

  • Native multimodal integration
  • Supports function calling for action-oriented tasks
  • Fast and responsive streaming
  • Strong alignment between visual and textual understanding
  • Versatile for a variety of "small" multimodal tasks

Limitations

  • Limited context window compared to text-only models
  • Not optimized for deep logical reasoning
  • Higher output cost than some basic vision models

Pricing

  • Input: ... per million tokens
  • Output: ... per million tokens

When to Use

Choose this model for modern, interactive applications where the AI needs to "see" and "talk" simultaneously and potentially trigger actions in other software via function calling.

When to Choose a Different Model

For heavy-duty document analysis, use Chat, Vision, Document Analysis & Reasoning - Medium. For pure reasoning tasks, use Reasoning & Problem Solving - Small.


Reasoning & Agent tasks - Large

Description

A powerhouse for developers building autonomous systems. This model is optimized for powerful reasoning, agentic tasks, and versatile developer use cases, allowing it to plan, execute, and refine complex workflows independently.

Specifications

  • Context window: 65,536 tokens
  • Max output tokens: Not specified
  • Vision support: No
  • Reasoning mode: Yes
  • Function calling: Yes
  • Streaming: Yes
  • Availability: API only

Ideal Use Cases

  • Autonomous AI agent development
  • Complex software engineering tasks
  • Multi-step business process automation
  • Advanced data analysis and synthesis
  • Tool-use orchestration
  • Complex logical planning and execution

Strengths

  • Advanced reasoning for agentic behavior
  • Robust function calling for external tool use
  • High reliability in multi-step task execution
  • Optimized for developer-centric workflows
  • Strong analytical capabilities

Limitations

  • No vision support
  • Higher cost than basic chat models
  • Not optimized for creative, long-form prose

Pricing

  • Input: ... per million tokens
  • Output: ... per million tokens

When to Use

Select this model when building AI agents that must operate autonomously, use tools, and reason through complex problems to reach a goal. It is the premier choice for "Agentic" AI.

When to Choose a Different Model

If your agent needs to process images or documents, use Chat, Document Analysis & Agent tasks - Xtra Large. For simple chat, Chat, Multi-lingual, Coding & function calling - Small is more efficient.


Reasoning & Problem Solving - Medium

Description

A mid-tier reasoning model that provides a significant boost in logical depth over the Small version. It is optimized for thinking and reasoning, making it suitable for professional-grade analytical tasks.

Specifications

  • Context window: 32,768 tokens
  • Max output tokens: Not specified
  • Vision support: No
  • Reasoning mode: Yes
  • Function calling: Yes
  • Streaming: Yes
  • Availability: API only

Ideal Use Cases

  • Professional financial analysis
  • Complex logical auditing
  • Mid-level software architecture planning
  • Detailed technical troubleshooting
  • Advanced mathematical reasoning
  • Strategic planning assistance

Strengths

  • Stronger logical deduction than Small reasoning models
  • Full function calling support
  • Reliable step-by-step thinking
  • Balanced speed and depth
  • High accuracy on complex logical queries

Limitations

  • No vision support
  • Limited context window (32K)
  • Higher cost than the Small reasoning model

Pricing

  • Input: ... per million tokens
  • Output: ... per million tokens

When to Use

Use this model for professional analytical tasks where accuracy and logical rigor are paramount, but the extreme scale of the Xtra Large model is not required.

When to Choose a Different Model

For the highest possible reasoning performance, use Reasoning & Problem Solving - Xtra Large. For agentic workflows, use Reasoning & Agent tasks - Large.


Reasoning & Problem Solving - Xtra Large

Description

The pinnacle of logical deduction in the portfolio. This model is optimized for the most demanding reasoning chat completions, capable of handling abstract problems and complex logical chains with extreme precision.

Specifications

  • Context window: Not specified
  • Max output tokens: Not specified
  • Vision support: No
  • Reasoning mode: No
  • Function calling: No
  • Streaming: Yes
  • Availability: API only

Ideal Use Cases

  • Advanced scientific research analysis
  • Complex legal reasoning and case analysis
  • High-level mathematical proofs
  • Deep architectural software design
  • Strategic corporate planning
  • Complex logic-based auditing

Strengths

  • Highest level of logical reasoning available
  • Capable of handling extremely abstract problems
  • High precision in complex deductions
  • Large-scale reasoning capacity
  • Optimized for complex chat completions

Limitations

  • Highest latency among reasoning models
  • Premium pricing
  • No vision support

Pricing

  • Input: ... per million tokens
  • Output: ... per million tokens

When to Use

Deploy this model for mission-critical analytical tasks where a failure in logic is unacceptable and the complexity of the problem requires the maximum available "thinking" capacity.

When to Choose a Different Model

If you need vision capabilities alongside reasoning, use Chat, Document Analysis & Agent tasks - Xtra Large. For faster, simpler logic, use Reasoning & Problem Solving - Small.


Chat & Function Calling - Small (Granite 3.1)

Description

Based on the IBM Granite 3.1 8B Instruct architecture, this is a long-context model optimized for instruction following, RAG, and function calling. It is highly efficient and supports 12 languages, including English, German, French, Italian, and Dutch.

Specifications

  • Context window: 131,072 tokens
  • Max output tokens: Not specified
  • Vision support: No
  • Reasoning mode: No
  • Function calling: Yes
  • Streaming: Yes
  • Availability: API only

Ideal Use Cases

  • RAG (Retrieval Augmented Generation) pipelines
  • Multilingual text extraction and summarization
  • API-driven automation
  • High-volume instruction following
  • Multilingual customer service bots
  • Structured data generation

Strengths

  • Excellent long-context handling (131K)
  • Strong function calling capabilities
  • Optimized for RAG workflows
  • Broad European language support
  • Very cost-effective

Limitations

  • No vision support
  • No dedicated reasoning mode
  • Not designed for complex creative writing

Pricing

  • Input: ... per million tokens
  • Output: ... per million tokens

When to Use

Use this model for RAG applications or any workflow that requires processing large amounts of text and then calling a function to act on that data, especially in a multilingual European context.

When to Choose a Different Model

For tasks requiring deep reasoning, use Reasoning & Agent tasks - Large. For vision tasks, use Chat, Vision, Document Analysis & Reasoning - Medium.


Reasoning & Tool Use - Large (GLM-4.5 Air)

Description

A Mixture-of-Experts (MoE) model featuring 106B total parameters. It offers hybrid reasoning with a configurable thinking mode, strong tool/function calling, and exceptional code generation capabilities, all within a generous context window.

Specifications

  • Context window: 131,072 tokens
  • Max output tokens: Not specified
  • Vision support: No
  • Reasoning mode: Yes
  • Function calling: Yes
  • Streaming: Yes
  • Availability: API only

Ideal Use Cases

  • Advanced code generation and refactoring
  • Complex tool-use orchestration
  • Hybrid reasoning tasks (fast vs. deep)
  • Large-scale technical documentation analysis
  • Developer productivity tools
  • Complex API integration workflows

Strengths

  • Efficient MoE architecture
  • Configurable thinking mode for flexibility
  • Strong coding and technical capabilities
  • Large 128K context window
  • Robust function calling

Limitations

  • No vision support
  • Higher output cost than basic chat models
  • Complexity in configuring thinking modes

Pricing

  • Input: ... per million tokens
  • Output: ... per million tokens

When to Use

Choose this model for technical and developer-centric tasks, particularly those involving coding or complex tool orchestration where a large context window and flexible reasoning are required.

When to Choose a Different Model

For pure reasoning without the MoE complexity, use Reasoning & Problem Solving - Medium. For vision-based agent tasks, use Chat, Document Analysis & Agent tasks - Xtra Large.


Search, Chat & Analysis - Small

Description

A multimodal model optimized for web search and conversational AI. It is particularly suited for creative professionals, artists, and content creators who need a blend of search capabilities, visual understanding, and storytelling fluency.

Specifications

  • Context window: Not specified
  • Max output tokens: Not specified
  • Vision support: Yes
  • Reasoning mode: No
  • Function calling: No
  • Streaming: Yes
  • Availability: API only

Ideal Use Cases

  • Creative storytelling and narrative design
  • Web-based research and synthesis
  • Visual content inspiration and brainstorming
  • Artistic project planning
  • Content creation for marketing
  • Image-based research queries

Strengths

  • Integrated web search capabilities
  • Vision support for visual research
  • High creativity and fluency in prose
  • Fast streaming for interactive sessions
  • Versatile for non-technical creative work

Limitations

  • No function calling support
  • No dedicated reasoning mode
  • Context window not specified

Pricing

  • Input: ... per million tokens
  • Output: ... per million tokens

When to Use

Choose this model for creative workflows, storytelling, or research tasks that require a combination of web access and visual understanding.

When to Choose a Different Model

For structured business automation, use Chat & Function Calling - Small (Granite 3.1). For deep logical analysis, use Reasoning & Problem Solving - Small.


Chat & Document Analysis & Reasoning - Large

Description

A large-scale model delivering frontier-level performance across a broad range of complex tasks. It combines advanced multilingual capabilities with a reasoning mode that can be enabled to dynamically tailor responses based on the complexity of the query.

Specifications

  • Context window: Not specified
  • Max output tokens: Not specified
  • Vision support: Yes
  • Reasoning mode: Yes
  • Function calling: Yes
  • Streaming: Yes
  • Availability: API only

Ideal Use Cases

  • Complex enterprise document analysis
  • High-stakes multilingual business communication
  • Advanced reasoning for business strategy
  • Visual document auditing
  • Complex content validation
  • Multimodal executive reporting

Strengths

  • Frontier-level performance on complex tasks
  • Dynamic reasoning mode
  • Integrated vision and function calling
  • Exceptional multilingual capabilities
  • Versatile across modalities

Limitations

  • Higher cost than medium-tier models
  • Context window not specified
  • Higher latency when reasoning mode is active

Pricing

  • Input: ... per million tokens
  • Output: ... per million tokens

When to Use

Deploy this model for high-complexity business tasks that require a blend of vision, reasoning, and multilingual fluency, where the highest possible quality of output is required.

When to Choose a Different Model

For autonomous agent workflows, use Chat, Document Analysis & Agent tasks - Xtra Large. For simple, fast chat, use Chat, Multi-lingual, Coding & function calling - Small.


Document Analysis & OCR - Small (DeepSeek OCR)

Description

A specialized 3B parameter vision-language model engineered specifically for optical character recognition (OCR) and document understanding. It excels at converting complex documents into structured text or markdown, including table extraction and mathematical notation.

Specifications

  • Context window: 8,192 tokens
  • Max output tokens: Not specified
  • Vision support: Yes
  • Reasoning mode: No
  • Function calling: No
  • Streaming: Yes
  • Availability: API only

Ideal Use Cases

  • Converting PDFs/Images to Markdown
  • Complex table extraction from documents
  • Mathematical formula recognition
  • Digitizing handwritten notes
  • Structured data extraction from forms
  • High-precision OCR for archives

Strengths

  • Specialized in OCR and document structure
  • Exceptional table and math recognition
  • High precision in text extraction
  • Efficient 3B parameter size
  • Fast streaming of extracted text

Limitations

  • Very small context window (8K)
  • Not designed for general chat or reasoning
  • No function calling support

Pricing

  • Input: ... per million tokens
  • Output: ... per million tokens

When to Use

Use this model exclusively for OCR and document digitization tasks where the goal is to turn a visual document into a structured text format with high precision.

When to Choose a Different Model

For general document Q&A or chat, use Chat, Vision, Document Analysis & Reasoning - Medium. For agentic workflows, use Chat, Document Analysis & Agent tasks - Xtra Large.


Chat, Multi-lingual, Coding & function calling - Small

Description

A versatile, high-efficiency model that balances chat fluency, multilingual support, and technical capabilities. It is particularly strong in coding tasks and function calling, making it a reliable choice for developer-centric chat applications.

Specifications

  • Context window: 128,000 tokens
  • Max output tokens: Not specified
  • Vision support: No
  • Reasoning mode: No
  • Function calling: Yes
  • Streaming: Yes
  • Availability: Chat UI & API

Ideal Use Cases

  • Coding assistance and snippet generation
  • Multilingual developer chatbots
  • API-integrated chat applications
  • Technical support automation
  • Structured text generation
  • Fast multilingual correspondence

Strengths

  • Strong coding capabilities
  • Full function calling support
  • Large 128K context window
  • Balanced performance across multiple domains
  • Available in both Chat UI and API

Limitations

  • No vision support
  • No dedicated reasoning mode
  • Not optimized for extremely deep logical deduction

Pricing

  • Input: ... per million tokens
  • Output: ... per million tokens

When to Use

Choose this model for general-purpose technical chat, coding help, or any application that requires a mix of multilingual fluency and the ability to call external functions.

When to Choose a Different Model

For deep reasoning tasks, use Reasoning & Problem Solving - Small. For vision tasks, use Chat, Vision, Document Analysis & Reasoning - Medium.


Chat, Document Analysis, Coding & Reasoning - Xtra Large

Description

A multimodal powerhouse optimized for the intersection of chat, document analysis, coding, and reasoning. This model is designed for high-complexity technical workflows that require both visual understanding and deep logical processing.

Specifications

  • Context window: 1,000,000 tokens
  • Max output tokens: Not specified
  • Vision support: Yes
  • Reasoning mode: Yes
  • Function calling: Yes
  • Streaming: Yes
  • Availability: Chat UI & API

Ideal Use Cases

  • Analysis of massive technical codebases
  • Complex multimodal data analysis
  • Long-form technical reasoning
  • Large-scale document auditing with vision
  • Advanced software architecture analysis
  • Comprehensive data synthesis from mixed sources

Strengths

  • Unprecedented 1M token context window
  • Full multimodal capabilities (Vision + Text)
  • Integrated reasoning and function calling
  • Strong coding and data analysis performance
  • Available in both Chat UI and API

Limitations

  • Higher cost per token
  • Higher latency for very large context windows
  • May be overkill for simple chat tasks

Pricing

  • Input: ... per million tokens
  • Output: ... per million tokens

When to Use

Deploy this model when you need to process an enormous amount of information (up to 1M tokens) that includes a mix of code, text, and images, and requires deep reasoning to synthesize the results.

When to Choose a Different Model

For faster, shorter interactions, use Chat, Vision, Document Analysis & Reasoning - Medium. For pure OCR, use Document Analysis & OCR - Small (DeepSeek OCR).


Chat, Vision, Document Analysis & Reasoning - Medium

Description

A best-in-class multimodal model that provides a high-performance balance of vision, coding, and reasoning. It is designed to be the versatile "go-to" model for most professional multimodal applications.

Specifications

  • Context window: 256,000 tokens
  • Max output tokens: Not specified
  • Vision support: Yes
  • Reasoning mode: Yes
  • Function calling: Yes
  • Streaming: Yes
  • Availability: Chat UI & API

Ideal Use Cases

  • Professional multimodal assistants
  • Technical document analysis with reasoning
  • Mid-to-large scale coding tasks
  • Visual data analysis and reporting
  • Complex multilingual chat with vision
  • General purpose high-end business AI

Strengths

  • Excellent balance of speed and capability
  • Strong multimodal and reasoning integration
  • Large 256K context window
  • Very competitive pricing for its capability tier
  • Available in both Chat UI and API

Limitations

  • Not as deep as Xtra Large for massive datasets
  • Higher cost than basic chat models
  • Reasoning can increase latency

Pricing

  • Input: ... per million tokens
  • Output: ... per million tokens

When to Use

Use this model as your primary multimodal engine for tasks that require a mix of vision, reasoning, and coding, where a 256K context window is sufficient.

When to Choose a Different Model

For the absolute maximum context (1M tokens), use Chat, Document Analysis, Coding & Reasoning - Xtra Large. For simple text chat, use Chat, Multi-lingual, Coding & function calling - Small.


Chat, Vision, Document Analysis, Coding & Reasoning - Medium (Gemma 4)

Description

A state-of-the-art multimodal model optimized for a comprehensive suite of capabilities including chat, vision, document analysis, coding, and reasoning. It represents the latest in efficient, high-performance AI for professional workflows.

Specifications

  • Context window: 256,000 tokens
  • Max output tokens: Not specified
  • Vision support: No
  • Reasoning mode: No
  • Function calling: No
  • Streaming: Yes
  • Availability: Chat UI & API

Ideal Use Cases

  • Advanced multimodal business automation
  • Technical coding assistance with document context
  • Complex visual data synthesis
  • High-performance multilingual chat
  • Professional document reasoning
  • Integrated coding and analysis workflows

Strengths

  • Latest generation multimodal architecture
  • Strong balance of coding and reasoning
  • Large 256K context window
  • High efficiency and response quality
  • Available in both Chat UI and API

Limitations

  • No vision support in this specific configuration
  • No function calling capabilities
  • Higher cost than basic small models

Pricing

  • Input: ... per million tokens
  • Output: ... per million tokens

When to Use

Choose this model when you need the latest generation of performance for coding, reasoning, and document analysis within a large context window.

When to Choose a Different Model

If you specifically require vision support, use Chat, Vision, Document Analysis & Reasoning - Medium. For the largest possible context, use Chat, Document Analysis, Coding & Reasoning - Xtra Large.


inference-miner-u25

Description

A specialized vision-language model strictly optimized for the technical tasks of document analysis and parsing. It is designed to extract structure and meaning from visual documents with high efficiency.

Specifications

  • Context window: Not specified
  • Max output tokens: Not specified
  • Vision support: No
  • Reasoning mode: No
  • Function calling: No
  • Streaming: Yes
  • Availability: API only

Ideal Use Cases

  • High-volume document parsing
  • Automated data extraction from forms
  • Visual structure analysis
  • Batch document processing
  • Industrial document digitization
  • Parsing of standardized business reports

Strengths

  • Highly optimized for parsing workflows
  • Efficient processing of visual layouts
  • Fast streaming for pipeline integration
  • Reliable for structured extraction
  • Cost-effective for parsing-specific tasks

Limitations

  • No general chat capabilities
  • No reasoning or function calling
  • No vision support (optimized for parsing)

Pricing

  • Input: ... per million tokens
  • Output: ... per million tokens

When to Use

Deploy this model within a backend pipeline specifically for the purpose of parsing documents and extracting data into a structured format.

When to Choose a Different Model

For any task requiring a conversational interface or reasoning, use Chat, Vision, Document Analysis & Reasoning - Medium. For high-precision OCR, use Document Analysis & OCR - Small (DeepSeek OCR).


Chat, Document Analysis & Agent tasks - Xtra Large

Description

Our most comprehensive model, designed for the most demanding enterprise applications. This very large-scale system combines vision, reasoning, and agentic capabilities with advanced multilingual support, enabling sophisticated automation across complex document workflows and integrated search operations.

Specifications

  • Context window: 250,000 tokens
  • Max output tokens: Not specified
  • Vision support: Yes
  • Reasoning mode: Yes
  • Function calling: Yes
  • Streaming: Yes
  • Availability: Chat UI & API

Ideal Use Cases

  • End-to-end enterprise automation pipelines
  • Complex agent orchestration with visual inputs
  • Vision-enabled document analysis and extraction
  • Advanced reasoning workflows with tool use
  • Multilingual agent deployment at scale
  • Integrated search, analysis, and action systems

Strengths

  • Comprehensive capability set (Vision, Reasoning, Function Calling)
  • Frontier-level performance across all modalities
  • Advanced multilingual support for global deployment
  • Full agentic functionality for autonomous operations
  • Enterprise-grade reliability

Limitations

  • Highest cost per token in the portfolio
  • Higher latency for complex reasoning tasks
  • May be overkill for simple chat

Pricing

  • Input: ... per million tokens
  • Output: ... per million tokens

When to Use

Select this model for mission-critical applications requiring the full spectrum of AI capabilities—vision, reasoning, and tool use—in a single integrated solution.

When to Choose a Different Model

For cost-sensitive applications not requiring all capabilities, consider Chat & Document Analysis & Reasoning - Large for vision without agentic focus, or Reasoning & Agent tasks - Large for reasoning without vision.


Reasoning & Tool Use - Xtra Large (GLM-5.1)

Description

The successor to the GLM-4.5 series, this 754B parameter model is a powerhouse of hybrid reasoning. It features a configurable thinking mode and exceptional capabilities in tool use, function calling, and complex code generation.

Specifications

  • Context window: Not specified
  • Max output tokens: Not specified
  • Vision support: No
  • Reasoning mode: Yes
  • Function calling: Yes
  • Streaming: Yes
  • Availability: Chat UI & API

Ideal Use Cases

  • Enterprise-grade software architecture
  • Complex tool-chain orchestration
  • High-precision technical reasoning
  • Advanced automated coding agents
  • Large-scale logical synthesis
  • Complex API-driven business logic

Strengths

  • Massive 754B parameter scale
  • Advanced hybrid reasoning capabilities
  • Top-tier function calling and tool use
  • Exceptional code generation
  • High reliability for complex tasks

Limitations

  • No vision support
  • Higher latency due to model scale
  • Premium pricing tier

Pricing

  • Input: ... per million tokens
  • Output: ... per million tokens

When to Use

Use this model when you need the absolute maximum in reasoning and tool-use capabilities for text-based tasks, particularly in software engineering or complex business logic automation.

When to Choose a Different Model

If you require vision capabilities, use Chat, Document Analysis & Agent tasks - Xtra Large. For faster, simpler reasoning, use Reasoning & Problem Solving - Medium.


Reasoning & Tool Use - Large (GLM-5)

Description

A high-performance reasoning model that brings the power of the GLM-5 architecture to a large-scale deployment. It offers hybrid reasoning with a configurable thinking mode and strong capabilities in function calling and code generation.

Specifications

  • Context window: Not specified
  • Max output tokens: Not specified
  • Vision support: No
  • Reasoning mode: Yes
  • Function calling: Yes
  • Streaming: Yes
  • Availability: API only

Ideal Use Cases

  • Technical tool orchestration
  • Complex code generation
  • Logical problem solving for developers
  • Automated API integration
  • High-precision technical chat
  • Complex reasoning pipelines

Strengths

  • Strong hybrid reasoning
  • Robust function calling
  • High-quality code output
  • Professional-grade logical deduction
  • Efficient for large-scale technical tasks

Limitations

  • No vision support
  • API only availability
  • Higher cost than basic chat models

Pricing

  • Input: ... per million tokens
  • Output: ... per million tokens

When to Use

Deploy this model for technical workflows requiring strong reasoning and tool integration where API access is the primary interface.

When to Choose a Different Model

For Chat UI access and even greater scale, use Reasoning & Tool Use - Xtra Large (GLM-5.1). For vision-enabled reasoning, use Chat, Document Analysis & Agent tasks - Xtra Large.


inference-multilingual-e5-large — intfloat/multilingual-e5-large

Description

A specialized embedding model designed to convert text into high-dimensional vectors. It is used primarily for semantic search, retrieval-augmented generation (RAG), and clustering across multiple languages.

Specifications

  • Context window: Not specified
  • Max output tokens: Not specified
  • Vision support: No
  • Reasoning mode: No
  • Function calling: No
  • Streaming: No
  • Availability: API only

Ideal Use Cases

  • Multilingual semantic search
  • RAG pipeline vectorization
  • Document clustering and grouping
  • Cross-lingual information retrieval
  • Text similarity analysis
  • Knowledge base indexing

Strengths

  • High-quality multilingual embeddings
  • Optimized for retrieval tasks
  • Consistent vector space across languages
  • Efficient for indexing large datasets

Limitations

  • Not a generative model (cannot "chat")
  • No reasoning or vision capabilities
  • API only

Pricing

  • Input: Not specified
  • Output: Not specified

When to Use

Use this model as the embedding engine for your search or RAG infrastructure when you need to support multiple languages with high semantic accuracy.

When to Choose a Different Model

This is a non-generative model. For generating responses based on retrieved data, pair this with Chat & Function Calling - Small (Granite 3.1).


Decommissioned Models

Chat & Document Analysis - Medium — decommissioned on 2026-06-01, replaced by Chat, Multi-lingual, Coding & function calling - Small Chat & Document Analysis & Reasoning - Large — decommissioned on 2026-06-01, replaced by Chat, Document Analysis, Coding & Reasoning - Xtra Large Search, Chat & Analysis - Large — decommissioned on 2026-06-01, replaced by Chat, Document Analysis, Coding & Reasoning - Xtra Large Apertus Swiss LLM - Small — decommissioned on 2026-05-28, replaced by Apertus Swiss LLM - Large Document Analysis - Medium — decommissioned on 2026-05-28, replaced by Chat & Document Analysis & Reasoning - Large Llama 3.3 Multi-lingual - Medium — decommissioned on 2026-05-28, replaced by Llama 4 Maverick multi modal - Small Reasoning & Problem Solving - Small — decommissioned on 2026-05-28, replaced by Reasoning & Problem Solving - Medium Reasoning & Problem Solving - Xtra Large — decommissioned on 2026-05-28, replaced by Reasoning & Problem Solving - Xtra Large Chat & Vision - Small (Gemma 3n) — decommissioned on 2026-06-01, replaced by Chat, Vision, Document Analysis, Coding & Reasoning - Medium (Gemma 4) Reasoning & Agentic - Large (GPT-OSS 120B) — decommissioned on 2026-05-11, replaced by Chat, Document Analysis, Coding & Reasoning - Xtra Large Embedding - Multilingual (BGE-M3) — decommissioned Embedding - Multilingual (Granite 278M) — decommissioned Chat & Document Analysis - Xtra Xtra Large — decommissioned on 2026-05-28 Chat, Document Analysis & Agent tasks - Xtra Large — decommissioned inference-bge-reranker — decommissioned

Model Updates

Schatzi AI continuously updates our model library to provide the latest frontier-level performance. Specifications, context windows, and pricing may be adjusted to reflect infrastructure improvements. For the most current comparison, please visit /ai-models/model-comparison.

Pricing Notice

Pricing is subject to change at our discretion.