AI on Demandpowered by Google

Gemma-4-31B-it — a multimodal open-weights model from Google DeepMind for text and image understanding, reasoning, coding and agentic workflows, operated by stepping stone entirely on Swiss infrastructure.

Gemma-4-31B-it is an instruction-tuned, multimodal open-weights model from Google DeepMind. The dense model has 30.7 billion parameters, processes text and images, and generates text output. Its context window of up to 160,000 tokens enables it to work with extensive documents, long conversations and larger codebases in context.

The model supports a configurable thinking mode for demanding reasoning as well as native function calling for structured integration with external tools. Gemma 4 is designed for text generation, visual understanding, coding, multilingual communication and agentic applications. Google states support for more than 35 languages and pre-training across more than 140 languages.

stepping stone operates gemma-4-31B-it entirely on Swiss infrastructure. Customers access the model through an OpenAI-compatible API and can integrate it directly into existing applications, knowledge systems, development environments and agentic workflows.

Gemma-4-31B-it is suitable for organisations that want to process text and visual content with a capable model without transferring their data to proprietary US-based AI platforms.

Typical use cases include analysing extensive documents, interpreting images and user interfaces, multilingual assistance, software development, structured information extraction and agentic workflows.

Open weights. Multimodal processing. Swiss infrastructure.

Gemma-4-31B-it combines text and image understanding with a large context window, configurable reasoning and native function-calling support. This allows a broad range of document-based and technical applications to be covered by a single model.

Operation by stepping stone keeps data processing on Swiss infrastructure and provides customers with a standardised API, predictable operation and personal support from our team in Bern.

Scope of services

AI model on demand

Access to gemma-4-31B-it for text and image understanding, analysis, reasoning, coding and agentic workflows through an OpenAI-compatible API.

GPU performance on demand

Scalable GPU infrastructure in Swiss data centres for demanding inference workloads, without the need to build and operate dedicated AI infrastructure internally.

Managed service

Deployment, monitoring, maintenance and support by stepping stone. We operate the model and infrastructure so customers can focus on integrating AI into their applications and workflows.

Areas of application

Document and image analysis

Gemma-4-31B-it processes text and images together, making it suitable for multimodal analysis tasks.

The model can analyse documents and screenshots, extract content through optical character recognition, and interpret charts, user interfaces and visual relationships. Text and image information can be combined in one request so the model considers both modalities together.

Analysis and reasoning

Gemma-4-31B-it supports demanding analysis and multi-step reasoning across long contexts.

Its configurable thinking mode allows applications to choose between direct response generation and additional computation for complex problems. The context window of up to 160,000 tokens supports extensive source material, longer conversations and connected information sets.

Software development and agentic workflows

Gemma-4-31B-it combines coding capabilities with native function-calling support.

The model can generate, complete, analyse and correct code. Combined with an agent harness and suitable tools, it can call external functions, evaluate intermediate results and work through tasks over several steps.

Benchmark

The benchmarks are performed with the vLLM serving benchmark. For a defined number of input and output tokens, the test measures metrics including throughput and end-to-end latency. Because results depend on prompt length, output length, concurrency and system load, different configurations can be tested and compared. Step-by-step instructions and the required benchmark scripts can be downloaded from GitHub.

 

Call

# Set your personal key:
STONEY_KEY=sk-...

# Make key visible for vllm bench:
export OPENAI_API_KEY=$STONEY_KEY

# Start the benchmark
vllm bench serve \
 --backend openai-chat \
 --model "google/gemma-4-31B-it" \
 --base-url llm.stoney-cloud.com \
 --endpoint /v1/chat/completions \
 --dataset-name random \
 --random-input-len 1024 \
 --random-output-len 256 \
 --num-prompts 50 \
 --max-concurrency 1 \
 --percentile-metrics e2el

 

Result

============ Serving Benchmark Result ============
Successful requests:                     50        
Failed requests:                         0         
Maximum request concurrency:             1         
Benchmark duration (s):                  231.27    
Total input tokens:                      51870     
Total generated tokens:                  12800     
Request throughput (req/s):              0.22      
Output token throughput (tok/s):         55.35     
Peak output token throughput (tok/s):    119.00    
Peak concurrent requests:                2.00      
Total token throughput (tok/s):          279.63    
----------------End-to-end Latency----------------
Mean E2EL (ms):                          4625.01   
Median E2EL (ms):                        4589.86   
P99 E2EL (ms):                           6077.83   
==================================================

 

Legend

  • Successful requests: Successful prompt requests.
  • Failed requests: Unsuccessful prompt requests.
  • Maximum request concurrency: Maximum number of requests processed simultaneously.
  • Benchmark duration (s): Duration of the benchmark run in seconds.
  • Total input tokens: Total number of input tokens.
  • Total generated tokens: Total number of tokens generated by the model.
  • Request throughput (req/s): Number of requests processed per second.
  • Output token throughput (tok/s): Average number of tokens generated per second.
  • Peak output token throughput (tok/s): Highest measured number of output tokens generated per second.
  • Peak concurrent requests: Highest measured number of requests processed simultaneously.
  • Total token throughput (tok/s): Average number of input and output tokens processed per second.
  • Mean End-to-End Latency (E2EL) (ms): Average time between a request and its completed output.
  • Median E2EL (ms): Median time between a request and its completed output.
  • p99 E2EL (ms): Latency not exceeded by 99 percent of successful requests.

Price

ModelContext lengthInput/MTokOutput/MTok
Gemma-4-31B-it160k0.25001.0000
All prices are in CHF/MTok, excluding VAT.

Wir verwenden Cookies, um unsere Website und unseren Service zu optimieren. Weitere Informationen findest du in unserer Datenschutzerklärung.