Logo
登录
Logo

最专注于开发者的语音AI平台

ISO 27001
ISO 27001
SOC 2
SOC 2
SSL/TLS
SSL/TLS
APPI
APPI
产品
  • 实时语音识别
  • 录音文件转写
  • 语音合成
  • 发音评测
  • DolphinTeams 双屏机
  • Tralingo AI翻译机
  • NihongoScore
资源
  • 文档
  • 博客
  • AI 应用
  • 在线体验
公司
  • 关于我们
  • 联系我们
  • 客户
法律
  • 隐私政策
  • 服务条款
  • 服务级别协议(SLA)
  • 基于特定商业交易法的标注
  • DolphinTeams 使用手册
© 2026 DolphinVoice All Rights Reserved.
产品LLMRTC Docs
LLMRTC Docs

LLMRTC Docs

OpenSource TypeScript SDK for real-time voice vision AI apps

0 点赞访问网站
LLMRTC Docs Screenshot
LLMRTC Docs Preview
访问网站

Introduction to LLMRTC Docs

LLMRTC is an open-source TypeScript SDK designed for building real-time voice and vision AI applications. It integrates WebRTC for low-latency audio and video streaming with large language models (LLMs), speech-to-text (STT), and text-to-speech (TTS) capabilities, all through a unified, provider-agnostic API. This allows developers to create seamless, interactive AI experiences without being locked into a specific cloud provider or infrastructure.

The SDK is structured into three main packages: @llmrtc/llmrtc-core for shared functionality, @llmrtc/llmrtc-backend for the Node.js server that handles WebRTC and provider orchestration, and @llmrtc/llmrtc-web-client for browser-based audio/video capture and playback. LLMRTC is ideal for developers looking to build voice assistants, customer support systems, multimodal agents, and on-device AI applications with minimal latency and high flexibility.

Takeaways

  • Real-time voice and vision capabilities with sub-second latency
  • Provider-agnostic architecture that supports multiple LLMs, STT, and TTS services
  • Tool calling and playbooks for multi-stage conversational flows
  • Streaming pipeline that improves user experience by starting responses before full generation
  • Session resilience with automatic reconnection and history preservation
  • Comprehensive observability with hooks and metrics tracking
  • Support for both cloud and local providers, enabling flexible deployment options

How LLMRTC Works

LLMRTC operates by combining WebRTC for real-time communication with AI model execution. On the backend, it manages audio and video streams, performs voice activity detection, and orchestrates interactions between different AI providers (e.g., using Claude for LLM, Whisper for STT, and ElevenLabs for TTS). The frontend SDK enables users to interact via voice, with real-time feedback and seamless integration of AI-generated responses.

The core workflow involves:

  1. User input (audio or visual) captured by the web client
  2. Streaming to the backend server for processing
  3. Execution of AI models and tools
  4. Generation of output (text or audio) sent back to the user
  5. Continuous interaction with hooks for logging, debugging, and custom behavior

Core Benefits and Applications

Use CaseDescription
Voice AssistantsBuild intelligent assistants with natural conversation flow and tool integration
Customer SupportImplement multi-step playbooks for efficient issue resolution
Multimodal AgentsCombine voice and vision for context-aware interactions
On-Device AIRun locally with no cloud dependencies for privacy and cost control

LLMRTC also provides extensive documentation, including quickstart guides, tutorials, and examples, making it accessible for developers at all levels.

标签

#TypeScript#WebRTC#LLM#Speech-to-Text#Text-to-Speech#AI SDK#Real-Time Streaming#Provider Agnostic#Multimodal AI#Developer Tools

精品推荐

Guideflow

Guideflow

The AI demo automation platform for SaaS

1259
CyberCut AI

CyberCut AI

AI video studio for viral social clips

706
Incredible

Incredible

Deep Work AI Agents - powered by Agent MAX

653
Typeless

Typeless

AI voice dictation that's actually intelligent

625

在 AI Apps 上免费展示您的应用

加入我们的创新者社区,让您的 AI 工具触达成千上万的每日用户。

申请展示
DolphinVoice Console

BlogPage.PromoContent.title

BlogPage.PromoContent.description

BlogPage.PromoContent.cta