Clera

ASR Engineer

San FranciscoFull TimePosted Sep 10, 2026

About the Role

This is an end-to-end ownership role for a cloud-based ASR and transcription pipeline at an early-stage ambient intelligence consumer startup. You'll work directly with product and general management leadership as one of the company's first US engineering hires, making real tradeoffs between latency, accuracy, and reliability as the product evolves.

What You'll Do

  • Build and iterate on the cloud-based ASR pipeline, from audio capture through post-processing, in production at scale.

  • Own ASR quality and reliability end-to-end, shipping measurable improvements on latency, small-word accuracy, and voice-print reliability.

  • Work across data, training and fine-tuning, evaluation, and deployment to turn product feedback into shipped pipeline changes.

  • Collaborate closely with overseas R&D, hardware, and supply-chain teams across time zones.

  • Partner with a product engineer on shared backend and pipeline surfaces.

  • Operate with minimal specification, translating lightweight asks into concrete, production-ready improvements.

What We're Looking For

  • 3+ years building and tuning transcription and ASR pipelines end-to-end in production, primarily in cloud-based settings.

  • Demonstrated ownership of production ASR systems through the full lifecycle: data preparation, model training and fine-tuning, evaluation, and deployment.

  • Hands-on experience with latency-sensitive or streaming audio and ASR pipelines.

  • Proficiency across the ML lifecycle, including data handling, evaluation metrics, and production deployment.

  • Track record of debugging and tuning transcription quality issues such as small-word accuracy, voice-print reliability, and latency.

  • Experience in early-stage or founding engineering environments, shipping without large team support or fully-specified requirements.

  • On-device or embedded ML experience (Core ML, TensorFlow Lite, or similar frameworks) is a plus.

  • Prior experience with wearable, hardware, or robotics products is a plus.

  • Background at AI-native consumer applications focused on transcription or audio is a plus.

  • Experience building agent or LLM-based product features, including tool use, memory, or retrieval, is a plus.

  • You care about how transcription feels to use, not just how it benchmarks, and can make latency and accuracy tradeoffs independently.

Compensation & Benefits

Salary range: $150,000 to $200,000 USD annually. Visa sponsorship is not available for this role.

Location

Hybrid, 3 days per week in office. San Francisco Bay Area, California, United States.

Keep exploring

View all software engineering jobs