Speech-to-text · Diarization · LLM speaker attribution

The speech API that knows who's talking

Re:WayAI turns real-world conversations into structured, speaker-attributed transcripts. State-of-the-art diarization, LLM speaker attribution and your own private library of enrolled voices — one API call away.

50 hours free transcription No credit card Audio deleted after processing
Accuracy
25% fewer errors

Relatively lower diarization error rate than pyannote community-1, the state-of-the-art open-source model, across four public benchmarks.

Languages
25

European languages in batch, 40 languages & regional variants in real-time streaming — same pipeline, same speaker accuracy.

Audio retained
0 bytes

Source audio is deleted the moment processing completes. Deleted, not archived.

To start
50 hours

Free speaker-attributed transcription on every new account. Per-minute billing after that.

How it works

From raw audio to named speakers

One job, four stages. Each stage returns its results as soon as it finishes — you never wait for the whole pipeline to read the transcript.

01 / INGEST

Upload or stream

Send any common audio or video format — conversion is handled. Or stream 16 kHz PCM over a WebSocket for real time.

02 / TRANSCRIBE

Accurate speech recognition

Word-level timestamps in 25 European languages, batch or streaming, built for conversational speech.

03 / DIARIZE

Who spoke when

State-of-the-art segmentation by voice — robust to overlap, crosstalk and distant microphones.

04 / LLM ATTRIBUTION

Real names

An LLM reads the conversation and names each speaker — with confidence and quoted evidence — or matches voices you've enrolled.

Accuracy

Measured, not promised

Diarization error rate (DER) across four public benchmark corpora — real meetings and in-the-wild media. Identical audio, identical scoring for every system.

Diarization error rate, drawn to scale bar length = DER · shorter is better · one shared scale
NOTSOFAR‑1 office meetings, single room microphone
pyannote community‑1
27.67%
Re:WayAI
20.43% −26% vs open source
ICSI natural research meetings, up to ~10 speakers
AWS Transcribe
46%
pyannote community‑1
30.84%
Re:WayAI
23.35% best
AMI project meetings, headset mix
Deepgram nova‑3
35%
AWS Transcribe
29%
pyannote community‑1
17.05%
pyannoteAI precision‑2
12.9%
Re:WayAI
12.85% best
VoxConverse in-the-wild media: debates, talk shows, interviews
Deepgram nova‑3
36%
AWS Transcribe
13%
pyannote community‑1
11.14%
pyannoteAI precision‑2
8.5%
Re:WayAI
8.62% 2nd best
Overall average across all four
pyannote community‑1
21.68%
Re:WayAI
16.31% −25% vs open source

DER — diarization error rate; best result per benchmark in bold. The green badge is our rank among the systems evaluated on that benchmark, or — where pyannote community‑1 is the only other system evaluated — our relative error reduction against it. Our numbers: internal evaluation, July 2026, identical audio and scoring for every system. precision‑2 as published in the pyannote‑audio benchmark; AWS and Deepgram via OpenBench. Bottom line: 25% relatively fewer diarization errors than the best open-source model.

cpWER is the strictest public metric for speaker-attributed transcription: a word counts as correct only if it is both transcribed correctly and attributed to the right speaker. Measured against nine commercial APIs on the same public benchmarks.

Speaker-attributed accuracy (cpWER), drawn to scale bar length = cpWER · shorter is better · one shared scale
NOTSOFAR‑1 dev office meetings, development set
Google
75.63%
Speechmatics
63.38%
Deepgram
62.55%
Mistral Voxtral
61.42%
Grok
58.92%
Gladia
56.12%
ElevenLabs
51.29%
AssemblyAI
48.22%
Azure
45.38%
Re:WayAI
47.89% 2nd best
NOTSOFAR‑1 test office meetings, single room microphone
Google
63.83%
Speechmatics
56.53%
Grok
52.35%
Mistral Voxtral
51.91%
Deepgram
48.24%
Gladia
47.81%
ElevenLabs
44.37%
AssemblyAI
37.02%
Azure
35.68%
Re:WayAI
43.03% 3rd best
AMI project meetings, headset mix
Google
46.96%
ElevenLabs
36.14%
Mistral Voxtral
34.86%
Gladia
34.34%
Grok
32.10%
Deepgram
29.28%
Azure
27.39%
AssemblyAI
27.36%
Speechmatics
20.82%
Re:WayAI
28.89% 4th best
DiPCo dinner parties, close-talk mix
Google
56.40%
ElevenLabs
48.78%
Gladia
41.90%
Grok
41.60%
Deepgram
38.31%
Mistral Voxtral
38.04%
Speechmatics
36.88%
AssemblyAI
33.48%
Azure
33.23%
Re:WayAI
39.21% 6th best
Overall average across all four
Google
60.70%
Mistral Voxtral
46.56%
Grok
46.24%
ElevenLabs
45.14%
Gladia
45.04%
Deepgram
44.60%
Speechmatics
44.40%
AssemblyAI
36.52%
Azure
35.42%
Re:WayAI
39.75% 3rd best

cpWER — best result per benchmark in bold. Competitor numbers as published on AssemblyAI’s benchmarks page (July 2026; Soniox omitted). Re:WayAI: internal evaluation on the same public corpora, same Whisper-normalizer protocol. Bottom line: third of ten systems overall — behind only Azure and AssemblyAI.

Capabilities

Built for the hard part: the speakers

Anyone can transcribe. We also get the speakers right.

Overlap-robust diarization

Segmentation that holds up on interruptions, crosstalk and distant-microphone recordings — not just clean studio audio.

LLM speaker attribution

An LLM reads the conversation and puts real names on speakers — each with a confidence score and the quoted evidence behind it.

Advanced LLM capabilities

The LLM stage goes beyond naming: it repairs diarization slips, cleans segment boundaries and turns raw output into a readable transcript.

Real-time streaming

Live transcription over a WebSocket with partial results as words are spoken — straight from a microphone or a call.

Transcript Q&A

Ask questions about any finished transcript — “what did we decide on pricing?” — and the answer is grounded strictly in what was said.

25 European languages

English plus 24 more, through the same pipeline — same diarization, same attribution, same accuracy.

Transcript editor

Fix a word, rename a speaker, pin a note to any segment — with the source video playing right next to the transcript.

One-click exports

Hand a clean DOCX or PDF to the humans — and JSON with word-level timings to the machines. Same job, both worlds.

Voice enrollment

Your own database of voices

Enroll a voice once with ~10 seconds of audio. From then on, that person is recognized by name in every recording you send — before the LLM even has to guess.

  • Private by design — your voice library belongs to your account alone; it is never shared or matched against anyone else's audio
  • Messy samples welcome — enrollment audio is diarized first, and the dominant speaker becomes the profile
Build your voice library
Privacy & security

Your audio never outlives the job

Voice recordings are some of the most sensitive data there is. The pipeline is built around that fact — GDPR compliant and ISO 27001 certified, not audited into shape afterwards.

Audio deleted immediately

Source audio is deleted the moment processing completes. We never retain, reuse or train on your recordings — there is nothing left to leak.

ISO 27001, our own data center

Re:WayAI is ISO 27001 certified. Every job runs inside our own certified data center — your audio never leaves it, and no third-party cloud or AI API ever touches it.

Transcripts stay yours

Finished transcripts are stored for you until you delete them — review, edit, annotate, export to PDF or DOCX, or remove them at any time via the console or the API.

On-premise

Audio can't leave the building at all?

The entire pipeline — models included — can be installed on your own hardware and run fully offline, behind your firewall.

Talk to us
ISO 27001 certified GDPR compliant EU data processing Encrypted in transit No training on your data On-premise available
Pricing

Transparent pricing, pay as you go

We run our own hardware in our own ISO 27001 data center — no big-cloud markup baked into every minute. That's exactly why we can undercut the big transcription APIs.

Speaker-attributed transcript with speaker names

The full pipeline in one API call: transcription, speaker diarization and LLM speaker naming — one price, billed per minute of audio.

  • No subscriptions, no minimums — pay only for audio you process
  • Or run any stage on its own: every job is billed per stage
  • Live rates in the console, metered usage API included
Understand conversation
from $0.095/hour

transcript · speakers · names

See pricing plans
Component — run only what you needpriceper minute
Speech-to-text (batch)25 European languages$0.025/hour$0.00042/min
Speech-to-text (real-time streaming)40 languages & regional variants · same price as batch$0.025/hour$0.00042/min
Speaker diarizationwho spoke when, enrolled-voice recognition included$0.06/hour$0.001/min
LLM speaker naming & every other AI promptreal names from the conversation · transcript Q&A billed per audio hour, never per token$0.01/hour$0.00017/min
Growth
from $0.095/hour

all-in — the rates in the table above

  • 50% off Starter on every stage, committed volume
  • Monthly or annual commitment, invoiced
  • Priority support
  • Same privacy guarantees: GDPR, ISO 27001
Speak with our experts
Enterprise
Custom

fully negotiable — built around your workload

  • On-premise: the whole pipeline behind your firewall
  • Custom SLAs (guaranteed uptime & response times) and dedicated support
  • Unlimited scale, custom integrations
  • Zero data retention by design
Speak with our experts
Contact

Speak with our experts

Volume pricing, on-premise deployment, or just not sure which stages you need — tell us about your workload and we'll get back to you within one business day.

We'll reply within one business day — you'll get an email confirmation right away.

Message sent

Thanks — we'll get back to you within one business day.
A confirmation is on its way to your inbox.