AI on Demand powered by OpenAI

Whisper Large V3 by OpenAI: the world’s most widely used open-source model for automatic speech recognition, trained on over 5 million hours of audio data. Supports around 100 languages, and performs reliably with accents and technical jargon — powered by stepping stone on Swiss infrastructure.

With Whisper Large V3 and GPT-OSS-120B, stepping stone provides two powerful OpenAI models for different use cases on Swiss infrastructure.

Whisper Large V3 is specialized in automatic speech recognition. It reliably transcribes spoken language in around 100 languages and supports timestamps as well as direct translation of spoken content into English.

GPT-OSS-120B is a powerful open-weight language model for complex text generation, analysis, reasoning, coding, and agentic workflows. Thanks to its Mixture-of-Experts architecture, it combines high model capacity with efficient inference.

The OpenAI models are suitable for a wide range of AI workloads — from speech and transcription to complex text, coding, and agentic applications.

Whisper Large V3: Transcription of meetings, interviews, and customer conversations, subtitling, translation, and accessibility.

GPT-OSS-120B: Text generation, analysis, reasoning, coding, RAG, tool calling, and agentic workflows.

For tasks with higher logical complexity or multi-step workflows, we recommend GPT-OSS-120B as a powerful foundation.

Open weights. Swiss data centers. No processing by US-based providers.

Specialized models for different tasks: Whisper Large V3 for speech and GPT-OSS-120B for demanding text, reasoning, and agentic applications. OpenAI-compatible interfaces enable easy integration into existing systems. Personal consulting and managed operation by stepping stone in Bern.

Scope of services

AI models on demand

Access to Whisper Large V3 for speech recognition and translation, as well as GPT-OSS-120B for text generation, analysis, reasoning, coding, and agentic workflows.

GPU performance on demand

Scalable computing power for individual recordings or large audio archives. From single transcriptions to bulk processing — you pay as you go.

Managed service

Deployment, monitoring, maintenance and support on Swiss infrastructure, with personalised advice. stepping stone takes care of the day-to-day running so that you can focus on the benefits.

Areas of application

Transcription

Whisper Large V3 is the industry standard for automatic speech recognition — tried and tested in thousands of production environments worldwide.

Teams use it to transcribe meetings, interviews, customer conversations and telephone calls. With timestamps at word and sentence level, content can be referenced precisely and fed into downstream processes.

Subtitling & Translation

Whisper reliably transcribes across around 100 languages — even when there is background noise, accents or technical jargon.

Media producers and platform operators use it for subtitling, automatic translation of spoken content into English, and accessible speech-to-text solutions. All of this is hosted on Swiss infrastructure, without any data being transferred to US services.

Reasoning, Coding & Agents

GPT-OSS-120B is suitable for more complex tasks where analysis, reasoning, coding, or multiple processing steps come together.

Typical use cases include knowledge assistants, RAG, code analysis and generation, structured data processing, tool calls, and agentic workflows.

Benchmark

The benchmarks were measured using the vllm bench tool against the production API gateway. The standard input sizes were 1,024 tokens. This equates to around 2–3 book pages or 500–750 words.

Call

# Set your personal key:
STONEY_KEY=sk-...

# Make key visible for vllm bench:
export OPENAI_API_KEY=$STONEY_KEY

# Start the benchmark
vllm bench serve \
 --backend openai-chat \
 --model "openai/gpt-oss-120b" \
 --base-url llm.stoney-cloud.com \
 --endpoint /v1/chat/completions \
 --dataset-name random \
 --random-input-len 1024 \
 --random-output-len 256 \
 --num-prompts 50 \
 --max-concurrency 1 \ \
 --percentile-metrics e2el

Result

============ Serving Benchmark Result ============
Successful requests:                     50        
Failed requests:                         0         
Maximum request concurrency:             1         
Benchmark duration (s):                  75.64     
Total input tokens:                      51468     
Total generated tokens:                  3300      
Request throughput (req/s):              0.66      
Output token throughput (tok/s):         43.63     
Peak output token throughput (tok/s):    248.00    
Peak concurrent requests:                2.00      
Total token throughput (tok/s):          724.06    
----------------End-to-end Latency----------------
Mean E2EL (ms):                          1512.53   
Median E2EL (ms):                        1502.29   
P99 E2EL (ms):                           2889.83   
==================================================

 

Legend

  • Successful requests: Successful prompt requests
  • Failed requests: Unsuccessful prompts
  • Maximum request concurrency: How many requests the model processes simultaneously.
  • Benchmark duration (s): The duration of the benchmark run in seconds.
  • Total input tokens: The total number of input tokens.
  • Total generated tokens: The total number of tokens generated by the model.
  • Request throughput (req/s): The number of requests processed per second.
  • Output token throughput (tok/s): The average number of tokens generated per second.
  • Peak output token throughput (tok/s): The maximum measured number of output tokens per second.
  • Peak concurrent requests: The maximum measured number of requests processed simultaneously.
  • Total token throughput (tok/s): The average of all tokens processed during the measurement.
  • Mean End to End Latency (E2EL) (ms): The average time between input and completed output.
  • Median E2EL (ms): The typical time between input and completed output.
  • p99 E2EL (ms): The "worst case" time until completed output.

Price

ModelMTok
whisper-large-v30.0020

ModelContext lengthInput/MTokOutput/MTok
gpt-oss-120b128k0.40001.6000

 

All prices are in CHF/MTok, excluding VAT.

Wir verwenden Cookies, um unsere Website und unseren Service zu optimieren. Weitere Informationen findest du in unserer Datenschutzerklärung.