GPUs Around the World
Your request heads to a nearby GPU, so distance does not turn a quick thought into a loading screen.
Built to Stay Fast
The same GPU network handles enormous AI workloads every day. You get that speed the instant you stop talking.
Always Warm
The model is already awake when you hit Stop. No cold start. No spin-up. No staring at a loader.
Ready for the Rush
First user or ten-thousandth, the pipeline spreads the work and keeps moving.
Six quick hops. Audio in. Text out. Nothing left behind.
MIC
Captured
ENCODE
WebM/Opus
API
Direct
INFER
DeepInfra
RETURN
< 1.8s
DELETE
Permanent
01
Browser Capture
Audio is captured natively in your browser using the WebAudio API. No plugin, no extension, no download required. Works on every modern device.
02
Efficient Encoding
Audio is encoded in WebM/Opus format, a codec purpose-built for voice. This minimizes file size and upload time while preserving every phoneme accurately.
03
Direct Request Path
Audio is sent through Yapr's transcription API for the live request, then handed to inference without creating a separate storage object.
04
AI Inference
Your audio is sent to DeepInfra's dedicated AI inference endpoint. State-of-the-art speech models run on dedicated GPU hardware, no shared queuing, no cold start, no delay.
05
Instant Return
Transcribed text is returned directly to your browser via our API. The median round-trip time is under 1.8 seconds for recordings under 60 seconds.
06
No Audio Object Storage
When transcription completes, audio is deleted immediately. Yapr has no recordings database, no audio archive, and no hidden storage layer.
0.2%
word accuracy
0K hrs
training data
0+
languages
0-bit
AES encryption
0.9%
service uptime
0bytes
audio retained
Accents. Background noise. Fast talkers. Language switches. YAPR keeps turning real speech into sharp, readable text.
Native English speakers
99.4%
Non-native English speakers
98.8%
Technical vocabulary
98.1%
Noisy environments
97.2%
Code-switching (2 languages)
96.9%
01
No Audio Storage Layer
The system is architected without an audio storage layer. Audio enters only to generate text. No recordings database. No audio archive. No backup of audio files.
02
Immediate Deletion
The transcription pipeline is built so audio is deleted immediately after transcription. No archive. No recordings database. No retention layer.
03
TLS 1.3 In Transit
All data in transit uses TLS 1.3, the current gold standard in transport encryption. This covers your browser, our API, and our AI infrastructure.
04
AES-256 At Rest
Transcript text and account data are stored in AES-256-GCM encrypted database partitions with key rotation. The encryption layer is enforced at the infrastructure level, not application level.
05
Secure Authentication
Authentication is available via OAuth 2.0 (Google, GitHub), email with encrypted password hashing, or passkeys (WebAuthn). Your password is never stored in plain text. Your biometrics never leave your device.
06
Hardened Security Headers
Every response enforces HSTS, Content-Security-Policy, X-Frame-Options, and SameSite=Strict cookies, preventing XSS, clickjacking, and session hijacking by default.
07
Metadata Separation
The only data stored is account and usage metadata, plus transcript text when you enable history. History is off by default. Server-side audio is never retained. Your recordings are deleted immediately after transcription.
08
Your Data. Your Call.
Want a copy or want it gone? You can export or delete your data from Settings whenever you choose.