Inbox: every channel, one agent

v0.25 lands inbound messaging on five platforms at once — Telegram, Slack, Discord, email, WhatsApp. The hard part wasn't writing five adapters. It was finding the one shape that made all five behave like the same thing.

Chihab
Building Tutti AI · · 9 min read

An agent your users can’t reach is a demo. An agent your users can reach is a product.

For most of Tutti’s life, that gap was bigger than the framework itself. You could define an agent in twenty lines, give it Stripe access, gate the destructive tools behind HITL, deploy it with one command — and then to actually let a user message it, you had to write the inbound plumbing yourself. Polling Telegram. Hosting a webhook for WhatsApp. Wrangling Slack’s app-level token. Threading email replies with In-Reply-To. Five different ad-hoc systems, all of which had to share connections with the outbound voices the agent already used.

In v0.25 we shipped @tuttiai/inbox — five platform adapters under one orchestrator: Telegram, Slack, Discord, email, WhatsApp. Twitter follows once its voice’s shared client lands. This post is how we got there: what we tried, what we threw away, and why each platform looks the way it does.

The problem isn’t five adapters. It’s one orchestrator.

The naive shape is “five inbox packages, one for each platform.” That sounds modular. It isn’t. Five adapters means five slightly different identity models, five rate-limit implementations, five places to forget the allow-list, five log-redaction policies. Six months in, you have five subtly drifted code paths and any change you make has to be made five times.

The right shape is one orchestrator with a per-platform adapter interface, and every policy — allow-list, per-user rate limit, per-chat serial queue, error handling, identity, log redaction — applied uniformly above the adapter. The adapter’s job is reduced to: receive messages from the platform, normalise them to InboxMessage, ship the agent’s reply back. That’s it.

interface InboxAdapter {
  start(): Promise<void>
  stop(): Promise<void>
  subscribeMessage(handler: (msg: InboxMessage) => Promise<void>): void
  sendReply(chatId: string, text: string, raw?: unknown): Promise<void>
}

Every adapter — Telegram polling, Slack Socket Mode, Discord Gateway, IMAP IDLE, WhatsApp Cloud API webhooks — implements this. The orchestrator never has to know which.

A platform is a voice

The second decision was where the platform code lives. Two options were on the table:

  1. @tuttiai/inbox owns every platform client itself, and the existing voices (@tuttiai/slack, @tuttiai/discord, etc) keep their separate outbound clients.
  2. Each platform owns one client, exposed by its existing voice, and the inbox adapter borrows it.

Option 1 is simpler at first read. Option 2 wins immediately when you remember the constraint: most platforms allow exactly one connection per token. Discord’s Gateway API rejects two simultaneous bot sessions. Telegram’s getUpdates poll allows only one session per bot token. WhatsApp’s webhook listens on one HTTPS port. With option 1 you’d open two Discord connections — one for the voice’s outbound, one for the inbox’s inbound — and Discord would silently disconnect both.

So option 2 is the only correct shape. Each voice exports a token-keyed client wrapper:

const wrapper = SlackClientWrapper.forToken(botToken, factory, { appToken })

And the inbox adapter dynamic-imports its voice and consumes that wrapper. A score that uses @tuttiai/discord for outbound tools AND @tuttiai/inbox’s Discord adapter for inbound shares one Gateway connection. The voice’s reference count holds the connection open until the last consumer releases it.

This is the rule we ended up writing on the wall: a platform is a voice. Every third-party platform integration lives under voices/<name>/, owns its client wrapper, and exposes the same *ClientWrapper.forToken(token) factory. The inbox is a thin orchestrator that borrows from voices.

Five platforms, five different inbound mechanisms

Once the shape was right, the per-platform work got specific. Each platform has its own answer to “how do new messages reach you” — and we kept the framework’s surface uniform without papering over the mechanism, because the mechanism leaks through to operations.

PlatformHow inbound arrivesThe thing we couldn’t paper over
TelegramLong-poll getUpdatesOnly one polling session per token. Inbox + voice must share.
SlackSocket Mode WebSocketTwo tokens — bot (xoxb-…) for outbound REST, app (xapp-…) for the socket.
DiscordPersistent Gateway connectionDMs require the DirectMessages intent. Made it a default.
EmailIMAP IDLE (server-pushed)RFC 5322 threading: In-Reply-To + References chain.
WhatsAppPublic HTTPS webhook with HMAC signatureYou need a public tunnel. There’s no polling fallback.

Each row is a paragraph in the docs. We considered hiding the differences behind a uniform “all platforms feel the same” abstraction. We didn’t, because the differences leak to the operator: Slack needs a second token; WhatsApp needs Cloudflare Tunnel or ngrok in front of port 3848; email needs an app-specific password if the provider has 2FA. Pretending these are the same would have meant five subtly broken first-runs.

The 24-hour rule that makes WhatsApp different

WhatsApp’s most surprising constraint isn’t the webhook. It’s the 24-hour customer-service window: outside of 24 hours since the user’s last inbound, the Cloud API rejects free-form messages with error 131047. Only pre-approved Message Templates work for re-engagement.

Hiding this would have produced a class of bug where the agent silently fails to deliver replies after 24 hours. So the voice ships two tools instead of one:

The agent has to choose. That’s deliberate. A framework that abstracts this away is a framework that lies about what’s possible.

Threading email replies — the part nobody else gets right

Email is the trickiest of the five. Inbound is easy enough — IMAP IDLE pushes new messages, no polling. The hard part is outbound replies that thread.

An email reply that doesn’t set In-Reply-To and extend References shows up in Gmail / Outlook / Apple Mail as a brand new conversation, unrelated to whatever the user wrote. The user wonders why your bot is writing them out of nowhere. The bot wonders why every conversation is one turn long.

The adapter caches a small bounded LRU (default 1000 entries) keyed by inbound Message-ID. On send, it builds:

That’s three lines of header construction and one cache lookup. Without them the entire feature is broken. The fact that we had to write this code, instead of nodemailer or some upstream library having it as the obvious default, is a small indictment of how this part of the email stack has aged.

Defence in depth, applied uniformly above the adapter

Because every adapter funnels through one orchestrator, every safety policy lands once. Defaults you don’t have to think about:

These are the kind of defaults that are hard to add later. Allow-list is easy. The bounded queue is hard — every existing user implicitly depends on the unbounded behaviour the moment you ship without it. We picked the boring-correct defaults on day one.

The other thing the orchestrator owns is identity. Default InMemoryIdentityStore is a union-find: link("telegram:42", "slack:U7") makes both sides resolve to the same Tutti session. A user who connects on Telegram, then later authenticates on Slack, sees one continuous conversation regardless of which channel the next message arrives on.

For multi-process deployments, swap in a Postgres- or Redis-backed store. The orchestrator’s identity lookup is one call; the store is pluggable behind the same interface as SessionStore and UserMemoryStore. The pattern is consistent across the runtime.

What we deferred, and why

A few things didn’t make v0.25:

The version of this post that ships in v0.26 will tell you which of those we got right.

The dogfood

The marketing-agent example score in the repo wires every adapter through to one support agent. The same agent fields Telegram DMs, Slack channel mentions, Discord pings, threaded emails, and WhatsApp messages — and the existing marketing orchestrator + twitter / discord / content specialists keep working unchanged. One agent, every channel.

That was the point. An agent your users can reach on whichever platform their day already lives on, without you writing the plumbing. Try it on a real workload and tell me where the abstraction breaks: the inbox guide is the fastest path; the @tuttiai/inbox README has the full API.

Older post
Smart routing: picking the cheapest model that can still do the job
8 min · Engineering
Newer post
What an agent needs after it works
11 min · Engineering

Start conducting.

One install. Your first agent running in 60 seconds. No signup. No telemetry.