What it had to do
The firm writes grant applications and advises nonprofits. Its knowledge lived in years of Google Drive folders, email threads and call notes, spread across many clients. The team wanted to ask a question in plain English and get an answer from their own material, with the source named, without one client’s documents ever showing up in another client’s work.
It also had to fit how the team already works: in Slack, with Word documents as the finished output, and with sensitive files kept out entirely.
What we built
- Answers from the firm’s own files and email, with the source documents named.
- Private client channels. Questions asked in a client’s channel only search that client’s material.
- Memory per channel that survives restarts, including the files people drop into the channel.
- Date-ordered lookups for questions like "what came in this week?"
- Word documents in the firm’s house style, delivered back into Slack.
- Guardrails: it searches before it claims to know something, never says it did something it didn’t, and takes pricing and scope only from the firm’s own documents.
- Sensitive files stay out. A credentials spreadsheet, for example, is excluded in code by its file ID, so renaming it doesn’t let it through.
How it works
Every three minutes, a background worker checks the firm’s Drive and email for anything new or changed. It reads the text, splits it into passages, and stores them in a search index tagged by client. The assistant answers in Slack from that index.
- Sources: the firm’s Google Drive files, email with attachments, and call transcripts.
- Every 3 minutes, new and changed items are picked up.
- A reader extracts the text. Slides and documents are read directly; AI vision is used only for image-only pages and scanned PDFs.
- The text is split into passages and stored in a knowledge base, tagged by client, with sensitive files excluded by name and file ID.
- The assistant answers in Slack, in private channels for each client, with memory and guardrails.
- The results are answers that name their sources, and Word documents in the firm’s house style.
Built on Claude, Postgres with vector search, and a durable job queue, on US-hosted infrastructure (DigitalOcean, New York; database in AWS us-east-2, Ohio). Google Drive and Gmail are connected with read-only access.
Tested, not assumed
Client separation. On 29 July 2026 we queried the live index for one client: all 20 passages returned belonged to that client, and none to anyone else. None of the 115 files shared in project channels leaked into general search.
Reliability. In the eight days to 17 September 2026, 99.87% of about 128,700 background runs succeeded. Most of the failures were a single burst of Google rate-limit errors on 16 September, and the next scheduled checks picked the work back up.
Freshness. When we measured it on 4 September, a newly filed call transcript was searchable about three and a half minutes after it landed in Drive.
Problems we hit, and how we fixed them
Uploading several large files at once could crash the assistant.
File reading now runs two at a time with a size-aware guard. Files of 18, 12 and 6 MB then went through with no crashes.
Call transcripts were being filed, but dozens were never connected to the assistant.
We connected them: 52 previously unread transcripts became searchable in one day.
Asked about the latest call, the assistant searched by topic and could miss it.
It now has a date-ordered lookup, so "did Friday’s call come in?" is answered first time.
Slides made of images, and scanned PDFs, came through with no text.
Slides are read directly first, and AI vision fills in only the pages that are pure images.
What it costs to run
The assistant and its background worker run on a single server that costs about $24 a month. On top of that are the managed database and AI usage. The chat gateway is open-source and self-hosted, so there is no licence fee.