Production system Case study 48+ days live Zero outages

AI SMS Sales Assistant for a Home-Services Company

A human-in-the-loop AI assistant that drafts on-brand text replies inside the team's existing SMS platform, grounded in the company's own conversation history, so reps answer leads in seconds instead of minutes, without losing their voice.

48+days live, zero outages 65 to 70%lower daily run cost Five-figureannual plan savings ~7 wkskickoff to production

The problem

Speed-to-lead was leaking deals, one text at a time.

  • Sales reps get a flood of inbound customer texts across many market inboxes. The speed-to-lead window is short, and slow replies lose deals.
  • Every reply was typed from scratch. Quality and voice were inconsistent, and high-volume or after-hours periods meant missed or delayed responses.
  • Off-the-shelf AI reply tools sound generic. They do not know this company's pricing, promotions, financing options, or how this team actually talks.

The solution

AI drafts, a human always sends.

A browser-based AI assistant lives inside the rep's existing SMS tool. The moment a customer texts, it generates three ready-to-send reply drafts in the rep's own voice. The rep edits if needed and sends with one click.

Nothing is ever sent automatically. A person is always in the loop, and the drafting model is Claude. High-risk messages, such as safety, legal, billing disputes, or a request for a manager, are detected and escalated to a team channel instead of being auto-drafted.

The trust model Context in, three smart drafts out, a human reviews and sends. The customer experience stays personal, and the assistant never invents a price.

How it works

Eight capabilities behind every draft.

Described at the capability level: what each part does and why it matters.

RAGGrounded in the company's own data

Every draft is informed by the team's real past conversations, retrieved on the fly from a managed vector database, so replies match how this team actually writes and what it offers.

VoicesPer-department assistants

Separate, purpose-built assistants for each line of business, sales, service, electrical, and more, each grounded in its own history and operating playbook. Not a one-size-fits-all bot.

LatencyInstant, zero-wait drafts

A pre-generation layer prepares the reply the instant a message arrives, so the draft is waiting before the rep opens the chat. No ten-second spinner.

ContextCustomer-aware

The assistant references the customer's real quote and what was discussed on a recent phone call, so the text picks up where the last touch left off, and it never invents a price.

SellConsultative, not an order-taker

Drafts acknowledge the customer, answer the question, ask a smart follow-up, and push to a call when a deal needs a human touch.

EstimateRep-in-loop auto-estimates

From a conversation, the assistant can draft a real estimate. The rep enters or accepts AI-suggested details and creates it, with identity fields auto-filled from the lead record.

LearnAlways learning

A nightly process feeds new real conversations back into the system, including the replies reps actually sent, so it keeps sounding more like the team over time.

CostCost-controlled by design

A routing layer skips junk messages, uses a cheaper model for simple replies and the top model for substantive ones, and caches the heavy shared context to keep running cost low.

The phased build

Shipped in phases, each one answering something the team actually needed.

1

Proof of concept Live, used daily

A message comes in, the AI returns three drafts in the panel. Simple, but real, and in a rep's hands fast on live customers.

2

The redesign that made it stick A deliberate half-step

The first rep said the panel blocked his screen, so it was rebuilt to float, drag, and resize. That turned "I tried it" into "I use it every day."

3

Voice and efficiency

Grounded the drafts in thousands of the rep's real past messages, added caching to cut cost, a nightly learning sync, and an alert-the-team path for anything unfamiliar instead of guessing.

4

Facts and consultative tone

Real estimate prices, recent-call context, a consultative rewrite, and the instant pre-generation that removed the wait.

5

Auto-estimates and multi-department rollout

Rep-in-loop estimate creation, plus purpose-built assistants cloned per department, each in its own voice.

Results and impact

Measured on real traffic, not a demo.

48+
Days live in production, zero outages
65 to 70%
Lower daily running cost
Five-figure
Saved per year on the messaging plan
Thousands
Real past messages grounding the rep's voice
43%
Lower cost per draft from routing and caching
3
Drafts ready per inbound message, instant
0
Auto-sends. A human reviews every reply
~7 wks
From kickoff to a live, daily-used system
Before
Before
After
After

Daily running cost down ~65 to 70%, plus ~43% lower per draft. No loss of quality or speed.

Running cost down ~65 to 70%, measured on real traffic, with no loss of quality.
Production note

When a third-party memory database had a brief outage, it was caught and fully restored the same day, with no customer impact, and the system caught up on everything it missed.

Facts, not guesses

The assistant quotes the customer's real estimate and recent-call context, never a made-up price. Auto-drafted estimates are reviewed by the rep before anything is sent.

Scope vs effort

One operator, working with AI augmentation, delivered the scope a 5 to 6 person development team would typically take 4 to 6 months to ship.

What it demonstrates

The skills behind the system.

Human-in-the-loop AI product design Retrieval-augmented generation on a private corpus Multi-agent / per-persona systems Prompt engineering Safety and escalation guardrails Latency engineering: pre-generation + caching LLM cost optimization Browser-extension integration Workflow-automation integration Production operations: monitoring, incident response Continuous learning pipelines Turning a messy sales operation into a shipped, daily-used system

Outcome and what is next

Live, expanding, and the first of several.

The system is live and expanding from the first department to every line of business, with auto-drafted estimates and continuous learning. It is the first of several AI initiatives in the engagement.