All projects

phone-agent

Persian AI phone receptionist: answers a real call over Asterisk AudioSocket, STT → LLM → TTS in natural Persian, with an Android companion app.

phone-agent architecture: caller, call forward, Asterisk AudioSocket, agent.py, STT, LLM, TTS, plus the Android client

A Persian (Farsi) AI phone receptionist: it answers a real call, listens, replies in natural Persian speech, and keeps short-term context across turns.

Highlights

  • Voice loop on a real phone line: Asterisk AudioSocket in, 8 kHz PCM framing and turn-taking by hand, speech back on the call.
  • Persian first: STT pinned to fa, Persian system prompt and fallbacks, Persian TTS voice.
  • Two delivery paths: SIP/GSM through Asterisk, or an Android default-dialer app that auto-answers and talks to the same HTTP turn API.
  • Small and dependency-light: one Python file on stdlib sockets and http.server plus httpx; models swap by env var.
  • Runs offline for a demo: scripts/demo_call.py drives the real call loop with stubbed models, no API key needed.

Why it’s interesting

  • Real-time voice loop over a telephony protocol: Asterisk streams raw 8 kHz PCM over AudioSocket, and agent.py frames/parses it, detects end of speech with a simple energy VAD, then runs STT, LLM and TTS per turn (handle_call, energy_vad_end).
  • Audio plumbing done by hand: 8 kHz telephone audio is wrapped as WAV for STT, and 24 kHz TTS output (MP3 decoded to PCM) is resampled back to 8 kHz for the line (resample, _mp3_to_pcm, play_pcm24_as_8k).
  • Persian-first: STT is pinned to fa with a Persian context prompt, the system prompt and fallback phrases are Persian, and TTS uses a Persian language boost. check.py round-trips TTS into STT and asserts the Persian greeting word comes back.
  • A second delivery path for phones without a SIP trunk: a small HTTP API (/v1/turn, /v1/greeting) and an Android app (Kotlin) that auto-answers and talks to it, plus a usb_bridge.py experiment that records the call over adb.
  • Everything model-related goes through one OpenRouter key, so STT, LLM and TTS models are swappable with env vars.

Architecture

flowchart LR
    Caller((Caller)) --> PSTN[SIP / GSM] --> AST[Asterisk]
    AST -- AudioSocket 8k PCM --> AG[agent.py]
    AND[Android app] -- HTTP /v1/turn WAV --> AG
    AG --> STT[OpenRouter STT]
    STT --> LLM[OpenRouter LLM]
    LLM --> TTS[OpenRouter TTS]
    TTS --> AG
    AG -- audio reply --> AST
    AG -- audio/wav --> AND

agent.py socket listens for AudioSocket connections from Asterisk (see asterisk-extensions.conf.example). agent.py serve runs both the AudioSocket listener and the HTTP API used by the Android client. Per-call history keeps the last 8 messages.

Tech stack

Python 3.12+ (stdlib sockets and http.server, httpx), OpenRouter (speech-to-text, chat, text-to-speech), Asterisk AudioSocket, Android (Kotlin, Gradle), pm2 (ecosystem.config.cjs).

Key techniques

  • AudioSocket frame protocol (type byte, 16-bit length, payload): _read_frame, _send_frame in agent.py.
  • Energy-based voice activity detection for turn taking: energy_vad_end.
  • Stateless-ish HTTP turn API with a session header for clients: PhoneHttp in agent.py.
  • Android default-dialer / in-call service that auto-answers and plays TTS through the speaker: android/app/src/main/java/dev/sepehr/phoneagent/.
  • Call-forwarding setup for a GSM SIM: CALL_FORWARDING.md.

Getting started

Try the call loop offline first (no key, no Asterisk; STT/LLM/TTS stubbed with demo data):

uv sync
uv run python scripts/demo_call.py

Then with a real key:

cp .env.example .env        # set OPENROUTER_API_KEY
uv sync
uv run python agent.py greet                       # writes the greeting wav
uv run python agent.py reply 'سلام، پیک هستم'       # one text turn, prints/saves the reply
uv run python agent.py serve                       # AudioSocket :9092 + HTTP :9093

Point Asterisk at the host using asterisk-extensions.conf.example. For the Android client, see android/INSTALL.md and set the server URL in the app.

Tests

There is no unit test suite. uv run python scripts/demo_call.py is an offline check of the AudioSocket framing, VAD, resampling and HTTP turn API (it asserts frame sizes, sample rates and that turn 2 sees turn 1’s history). uv run python check.py is a live smoke test (TTS then STT, must contain “سلام”) and needs a valid API key.

Screenshots

The project has no web UI, so the images here are the architecture diagram at the top and real terminal output, not app screenshots.

Offline demo call output

Real output of scripts/demo_call.py: one AudioSocket call and two HTTP turns (demo data, models stubbed).

The HTML sources for both images are in docs/diagrams/.

License

MIT


Built by Sepehr Radmard · LinkedIn · GitHub · more projects on my profile