Skip to content
Yapr
FreeTechUse CasesHow It WorksHelp
Sign inTry for free
Yapr

The world's fastest, most accurate voice transcription. Free.

Yapr Inc. 131 Hazelton Avenue, Yorkville Toronto, ON M5R 2E4, Canada

Product

Open app

Explore

FreeTechnologyUse CasesHow It WorksHelp CenterCompareTools

Legal

PrivacyTermsCookies

© 2026 Yapr Inc. All rights reserved.

Made loud by Meridii

Technology

THE ENGINE.

You talk. GPUs do the heavy lifting. Clean text comes flying back.

100+ languages · under 2 seconds · audio disappears after transcription

Start Free
Infrastructure

POWERED
BY
DEEPINFRA

01

GPUs Around the World

Your request heads to a nearby GPU, so distance does not turn a quick thought into a loading screen.

02

Built to Stay Fast

The same GPU network handles enormous AI workloads every day. You get that speed the instant you stop talking.

03

Always Warm

The model is already awake when you hit Stop. No cold start. No spin-up. No staring at a loader.

04

Ready for the Rush

First user or ten-thousandth, the pipeline spreads the work and keeps moving.

Model

Big Speech Models. Tiny Wait.
Built for real voices.

Architecture

Transformer

Deep transformer-based encoder-decoder architecture, trained end-to-end on hundreds of thousands of hours of real-world multilingual audio.

Parameters

1.5B+

1.5 billion learned parameters, trained on 680,000 hours of multilingual audio, one of the largest speech training datasets ever assembled.

Languages

100+

Natively understands 100+ spoken languages. No configuration needed, language is auto-detected, even when you switch languages mid-sentence.

WER (English)

2.7%

Word Error Rate of 2.7% on standard benchmarks, approaching human-level transcription accuracy across accents, dialects, and ambient noise.

Pipeline

From voice to text
in under 2 seconds.

Six quick hops. Audio in. Text out. Nothing left behind.

MIC

Captured

ENCODE

WebM/Opus

API

Direct

INFER

DeepInfra

RETURN

< 1.8s

DELETE

Permanent

01

Browser Capture

Audio is captured natively in your browser using the WebAudio API. No plugin, no extension, no download required. Works on every modern device.

02

Efficient Encoding

Audio is encoded in WebM/Opus format, a codec purpose-built for voice. This minimizes file size and upload time while preserving every phoneme accurately.

03

Direct Request Path

Audio is sent through Yapr's transcription API for the live request, then handed to inference without creating a separate storage object.

04

AI Inference

Your audio is sent to DeepInfra's dedicated AI inference endpoint. State-of-the-art speech models run on dedicated GPU hardware, no shared queuing, no cold start, no delay.

05

Instant Return

Transcribed text is returned directly to your browser via our API. The median round-trip time is under 1.8 seconds for recordings under 60 seconds.

06

No Audio Object Storage

When transcription completes, audio is deleted immediately. Yapr has no recordings database, no audio archive, and no hidden storage layer.

0.2%

word accuracy

0K hrs

training data

0+

languages

0-bit

AES encryption

0.9%

service uptime

0bytes

audio retained

Accuracy

99.2%
Word
Accuracy.

Accents. Background noise. Fast talkers. Language switches. YAPR keeps turning real speech into sharp, readable text.

Native English speakers

99.4%

Non-native English speakers

98.8%

Technical vocabulary

98.1%

Noisy environments

97.2%

Code-switching (2 languages)

96.9%

Privacy Architecture

Your audio has one job.
Become text.
Then disappear.

01

No Audio Storage Layer

The system is architected without an audio storage layer. Audio enters only to generate text. No recordings database. No audio archive. No backup of audio files.

02

Immediate Deletion

The transcription pipeline is built so audio is deleted immediately after transcription. No archive. No recordings database. No retention layer.

03

TLS 1.3 In Transit

All data in transit uses TLS 1.3, the current gold standard in transport encryption. This covers your browser, our API, and our AI infrastructure.

04

AES-256 At Rest

Transcript text and account data are stored in AES-256-GCM encrypted database partitions with key rotation. The encryption layer is enforced at the infrastructure level, not application level.

05

Secure Authentication

Authentication is available via OAuth 2.0 (Google, GitHub), email with encrypted password hashing, or passkeys (WebAuthn). Your password is never stored in plain text. Your biometrics never leave your device.

06

Hardened Security Headers

Every response enforces HSTS, Content-Security-Policy, X-Frame-Options, and SameSite=Strict cookies, preventing XSS, clickjacking, and session hijacking by default.

07

Metadata Separation

The only data stored is account and usage metadata, plus transcript text when you enable history. History is off by default. Server-side audio is never retained. Your recordings are deleted immediately after transcription.

08

Your Data. Your Call.

Want a copy or want it gone? You can export or delete your data from Settings whenever you choose.

READY TO
START?

No card. No setup marathon. Just talk.

Start Free How It Works