The mission
Make every KaiCalls call route observable, recoverable, and dependable from carrier ingress through customer outcome.
What you will own
- Own SIP, PSTN, programmable voice, call control, phone-number lifecycle, webhooks, recordings, and failover.
- Build observability across call legs, provider receipts, application state, and customer-visible outcomes.
- Lead incidents and reduce recurrence through tests, runbooks, and architecture changes.
- Support porting, forwarding, carrier escalation, and emergency rollback.
What success looks like
- Service-level indicators exist for routing success, latency, completion, and fallback.
- Carrier and application failures produce actionable alerts.
- Incident reviews produce durable, verified fixes.
- Real inbound and outbound routes are tested before changes are declared proven.
What we are looking for
- Deep production ownership of programmable voice, SIP, or carrier systems.
- Strong distributed-systems debugging and incident leadership.
- Experience with idempotency, retries, failover, and observability.
- Clear communication under pressure.
Helpful, not mandatory
- Twilio and Telnyx experience.
- SIP and RTP diagnostics.
- Number porting, CNAM, and messaging-registration familiarity.
- Node, TypeScript, and Postgres.
Compensation
The base salary range is $140,000-$185,000 USD for a US-remote employee. Final compensation will reflect role level, location policy, experience, benefits, variable compensation where applicable, and equity. Contractor arrangements are scoped separately.
How we evaluate
We use the same core evidence for every candidate: comparable outcomes, ownership, customer judgment, role skill, communication, and learning velocity. The process normally includes a short screen, a paid and bounded work sample, structured interviews, and reference checks.
Example paid work sample: Diagnose a sanitized trace where the carrier leg connected but the business owner's phone never rang. Explain missing evidence, mitigation, permanent repair, and the customer update.
Your first 90 days
- First 30 days: Map the production call path, risks, and incident history.
- By 60 days: Establish service indicators, alerts, and top runbooks.
- By 90 days: Eliminate the leading incident class and run a failover exercise.
How to apply
Email kai@kaicalls.com with the role title in the subject line. Include your resume or LinkedIn profile.