Inbox: every channel, one agent
v0.25 lands inbound messaging on five platforms at once — Telegram, Slack, Discord, email, WhatsApp. The hard part wasn't writing five adapters. It was finding the one shape that made all five behave like the same thing.
An agent your users can’t reach is a demo. An agent your users can reach is a product.
For most of Tutti’s life, that gap was bigger than the framework itself. You could define an agent in twenty lines, give it Stripe access, gate the destructive tools behind HITL, deploy it with one command — and then to actually let a user message it, you had to write the inbound plumbing yourself. Polling Telegram. Hosting a webhook for WhatsApp. Wrangling Slack’s app-level token. Threading email replies with In-Reply-To. Five different ad-hoc systems, all of which had to share connections with the outbound voices the agent already used.
In v0.25 we shipped @tuttiai/inbox — five platform adapters under one orchestrator: Telegram, Slack, Discord, email, WhatsApp. Twitter follows once its voice’s shared client lands. This post is how we got there: what we tried, what we threw away, and why each platform looks the way it does.
The problem isn’t five adapters. It’s one orchestrator.
The naive shape is “five inbox packages, one for each platform.” That sounds modular. It isn’t. Five adapters means five slightly different identity models, five rate-limit implementations, five places to forget the allow-list, five log-redaction policies. Six months in, you have five subtly drifted code paths and any change you make has to be made five times.
The right shape is one orchestrator with a per-platform adapter interface, and every policy — allow-list, per-user rate limit, per-chat serial queue, error handling, identity, log redaction — applied uniformly above the adapter. The adapter’s job is reduced to: receive messages from the platform, normalise them to InboxMessage, ship the agent’s reply back. That’s it.
interface InboxAdapter {
start(): Promise<void>
stop(): Promise<void>
subscribeMessage(handler: (msg: InboxMessage) => Promise<void>): void
sendReply(chatId: string, text: string, raw?: unknown): Promise<void>
}
Every adapter — Telegram polling, Slack Socket Mode, Discord Gateway, IMAP IDLE, WhatsApp Cloud API webhooks — implements this. The orchestrator never has to know which.
A platform is a voice
The second decision was where the platform code lives. Two options were on the table:
@tuttiai/inboxowns every platform client itself, and the existing voices (@tuttiai/slack,@tuttiai/discord, etc) keep their separate outbound clients.- Each platform owns one client, exposed by its existing voice, and the inbox adapter borrows it.
Option 1 is simpler at first read. Option 2 wins immediately when you remember the constraint: most platforms allow exactly one connection per token. Discord’s Gateway API rejects two simultaneous bot sessions. Telegram’s getUpdates poll allows only one session per bot token. WhatsApp’s webhook listens on one HTTPS port. With option 1 you’d open two Discord connections — one for the voice’s outbound, one for the inbox’s inbound — and Discord would silently disconnect both.
So option 2 is the only correct shape. Each voice exports a token-keyed client wrapper:
const wrapper = SlackClientWrapper.forToken(botToken, factory, { appToken })
And the inbox adapter dynamic-imports its voice and consumes that wrapper. A score that uses @tuttiai/discord for outbound tools AND @tuttiai/inbox’s Discord adapter for inbound shares one Gateway connection. The voice’s reference count holds the connection open until the last consumer releases it.
This is the rule we ended up writing on the wall: a platform is a voice. Every third-party platform integration lives under voices/<name>/, owns its client wrapper, and exposes the same *ClientWrapper.forToken(token) factory. The inbox is a thin orchestrator that borrows from voices.
Five platforms, five different inbound mechanisms
Once the shape was right, the per-platform work got specific. Each platform has its own answer to “how do new messages reach you” — and we kept the framework’s surface uniform without papering over the mechanism, because the mechanism leaks through to operations.
| Platform | How inbound arrives | The thing we couldn’t paper over |
|---|---|---|
| Telegram | Long-poll getUpdates | Only one polling session per token. Inbox + voice must share. |
| Slack | Socket Mode WebSocket | Two tokens — bot (xoxb-…) for outbound REST, app (xapp-…) for the socket. |
| Discord | Persistent Gateway connection | DMs require the DirectMessages intent. Made it a default. |
| IMAP IDLE (server-pushed) | RFC 5322 threading: In-Reply-To + References chain. | |
| Public HTTPS webhook with HMAC signature | You need a public tunnel. There’s no polling fallback. |
Each row is a paragraph in the docs. We considered hiding the differences behind a uniform “all platforms feel the same” abstraction. We didn’t, because the differences leak to the operator: Slack needs a second token; WhatsApp needs Cloudflare Tunnel or ngrok in front of port 3848; email needs an app-specific password if the provider has 2FA. Pretending these are the same would have meant five subtly broken first-runs.
The 24-hour rule that makes WhatsApp different
WhatsApp’s most surprising constraint isn’t the webhook. It’s the 24-hour customer-service window: outside of 24 hours since the user’s last inbound, the Cloud API rejects free-form messages with error 131047. Only pre-approved Message Templates work for re-engagement.
Hiding this would have produced a class of bug where the agent silently fails to deliver replies after 24 hours. So the voice ships two tools instead of one:
send_text_message— free-form text, valid only inside the 24h window. On 131047 it surfaces a hint pointing at the other tool.send_template_message— pre-approved templates, registered + approved in the Meta App console.
The agent has to choose. That’s deliberate. A framework that abstracts this away is a framework that lies about what’s possible.
Threading email replies — the part nobody else gets right
Email is the trickiest of the five. Inbound is easy enough — IMAP IDLE pushes new messages, no polling. The hard part is outbound replies that thread.
An email reply that doesn’t set In-Reply-To and extend References shows up in Gmail / Outlook / Apple Mail as a brand new conversation, unrelated to whatever the user wrote. The user wonders why your bot is writing them out of nowhere. The bot wonders why every conversation is one turn long.
The adapter caches a small bounded LRU (default 1000 entries) keyed by inbound Message-ID. On send, it builds:
Subject: Re: <original>— idempotent (no doubleRe:if already prefixed).In-Reply-To: <inbound-message-id>References: <existing-chain> <inbound-message-id>
That’s three lines of header construction and one cache lookup. Without them the entire feature is broken. The fact that we had to write this code, instead of nodemailer or some upstream library having it as the obvious default, is a small indictment of how this part of the email stack has aged.
Defence in depth, applied uniformly above the adapter
Because every adapter funnels through one orchestrator, every safety policy lands once. Defaults you don’t have to think about:
- Allow-list is off by default — every sender accepted — but
inbox.allowedUsers.<platform>locks down to a list of platform user ids. The lookup is constant-time per message. - Per-user token-bucket rate limit — 30 msg / 60s with burst 10 by default. A public bot endpoint is a DoS surface; we’d rather drop a spammer than have a single user burn an agent’s daily budget.
- Per-chat serial queue — depth 10. Reply for message N must ship before message N+1 runs against the same chat. Excess depth is dropped, not buffered indefinitely. (The wrong default would be unbounded buffering; you’d discover the bug at 3am when a user retried 200 times.)
inbox:message_receivedcarriestext_length, not the message body. The text is not in the event. If you need it, subscribe to the adapter directly. This prevents accidental PII leakage into traces and downstream telemetry.SecretsManager.redactruns on inbound text by default for email and WhatsApp. Strings shaped like API keys are stripped before the agent ever sees them. Opt out per platform when the agent legitimately needs to handle credentials.
These are the kind of defaults that are hard to add later. Allow-list is easy. The bounded queue is hard — every existing user implicitly depends on the unbounded behaviour the moment you ship without it. We picked the boring-correct defaults on day one.
Identity links across platforms
The other thing the orchestrator owns is identity. Default InMemoryIdentityStore is a union-find: link("telegram:42", "slack:U7") makes both sides resolve to the same Tutti session. A user who connects on Telegram, then later authenticates on Slack, sees one continuous conversation regardless of which channel the next message arrives on.
For multi-process deployments, swap in a Postgres- or Redis-backed store. The orchestrator’s identity lookup is one call; the store is pluggable behind the same interface as SessionStore and UserMemoryStore. The pattern is consistent across the runtime.
What we deferred, and why
A few things didn’t make v0.25:
- Twitter inbound — the voice exists for outbound, but its client wrapper hasn’t been ported to the shared-cache pattern yet. Until that lands, an inbox adapter would have to open a second client, which violates the “platform is a voice” rule. Better to wait.
- WhatsApp group chats — Meta’s Cloud API doesn’t deliver group messages over webhooks for two-way bots. There’s no workaround that doesn’t require unofficial WhatsApp Web automation, which violates ToS. Direct messages only.
- Outbound media — text only across all five platforms in v0.25. Inbound media (images, audio, files) surface as a placeholder + the resolved URL on
InboxMessage.raw. Adding a typedattachmentsfield to the canonicalInboxMessageshape is a deliberate cross-platform design pass — the right model needs all five platforms’ attachment semantics on the table at once. Deferred to v0.26.
The version of this post that ships in v0.26 will tell you which of those we got right.
The dogfood
The marketing-agent example score in the repo wires every adapter through to one support agent. The same agent fields Telegram DMs, Slack channel mentions, Discord pings, threaded emails, and WhatsApp messages — and the existing marketing orchestrator + twitter / discord / content specialists keep working unchanged. One agent, every channel.
That was the point. An agent your users can reach on whichever platform their day already lives on, without you writing the plumbing. Try it on a real workload and tell me where the abstraction breaks: the inbox guide is the fastest path; the @tuttiai/inbox README has the full API.