What an agent needs after it works

Every framework can produce a working demo. The hard part is everything after — running on a schedule, getting better at the same task, and not billing you while idle. Three operational primitives that close the gap, and why the framework — not the user — should own them.

Chihab
Building Tutti AI · · 11 min read

Every agent framework has the same demo. Twenty lines of TypeScript, an LLM, a tool that calls a real API. The REPL prints something useful. You screenshot it.

Most agents stop there. Not because the capabilities are missing — capabilities are the easy part. The hard part is everything after the demo. The agent only runs when you remember to run it. It does the same dance every time, never any faster. Its output lands in your terminal, not in the channel your team actually reads. And the deploy bill arrives whether anyone used it or not.

This post is about three things that close that gap. None of them are new research — they’re the boring operational glue that turns an agent from a parlour trick into something a team depends on. We’ve been building each one into Tutti because every team we talked to was reinventing the same three pieces, badly, and a framework that ducks them is a framework that stops at the demo.

The thread connecting them: the framework owns the operations, not the user. An agent shouldn’t have to know about cron timers, hibernation contracts, or which Slack token to use to post a daily summary. The score declares intent; the runtime makes it happen.


The problem: an agent that works isn’t an agent that ships

Imagine you’ve built a useful agent. A reviewer that reads PRs and posts comments. A reporter that pulls metrics from Stripe and writes a weekly summary. A standup bot that looks at GitHub activity and produces a “yesterday” digest.

Each works in isolation. You ran it from your terminal, it did what it said. The screenshot is great. Now ship it.

Three things tend to bite, in this order:

  1. It only runs when you run it. A working agent locked behind your terminal is a glorified script. The reporter needs to fire every Monday morning, the standup bot every weekday at 9, the reviewer when a PR opens. You can wire each one into GitHub Actions or a Heroku scheduler if you want — but every team rewrites that glue, and most teams get the credentials, the retry semantics, and the failure handling subtly wrong.

  2. Even when it runs, its output goes nowhere. The agent prints a markdown summary to stdout. Your team reads Slack. You write a wrapper that calls chat.postMessage, then a different wrapper for email, then a different one for Discord. Half of the agent’s system prompt becomes “format the output for Slack with the following rules…” — formatting concerns leak into reasoning concerns, and the prompt becomes a smell.

  3. The bill arrives whether or not anyone used it. You deployed the agent on Fly or Railway or a Kubernetes pod, and now it’s awake 24/7. Your real usage shape is bursty — five minutes of activity around 9am, silence for the next 23 hours — but you’re paying for 24 of them. Scale-to-zero exists. You can’t use it, because the agent uses an in-memory session store and an in-process scheduler, both of which evaporate the moment the function hibernates.

And one slower problem, the one teams hit a month in:

  1. It does the same work every time, the same slow way. The PR reviewer fetches the PR, fetches each changed file, reads them, posts a comment. Run after run after run. The agent doesn’t know it’s doing the same thing — every plan is rebuilt from scratch. Token-by-token, it pays the cognitive tax of re-deciding what it already decided last Tuesday.

These four problems look unrelated. They share a shape: operational concerns that don’t belong in agent code, but most frameworks force you to put them there.


How Tutti tackles each one

The framework’s job is to make each operational concern a config change, not a refactor. Three primitives, three corresponding score-level surfaces.

1. Scheduled delivery: the agent doesn’t post; the scheduler does

A schedule block on an agent declares when it runs, what input to seed it with, and (the new bit) where its reply should land:

import { defineScore, AnthropicProvider } from "@tuttiai/core";
import { GitHubVoice } from "@tuttiai/github";
import { SlackVoice } from "@tuttiai/slack";

export default defineScore({
  provider: new AnthropicProvider(),
  agents: {
    "standup-bot": {
      name: "Standup Bot",
      model: "claude-sonnet-4-6",
      voices: [new GitHubVoice(), new SlackVoice()],
      permissions: ["network"],
      schedule: {
        cron: "0 9 * * 1-5",
        input: "Summarise yesterday's PRs and issues across the team.",
        deliver: { platform: "slack", channel: "#standup" },
        deliver_format: "markdown",
      },
      system_prompt:
        "Read-only GitHub summary of yesterday. Markdown out. Do NOT post directly — the scheduler delivers.",
    },
  },
});

The agent’s only job is to produce the text. The scheduler dispatches the final reply to the configured target — Slack, Discord, Telegram, email, or WhatsApp — using the matching voice’s already-loaded client. Credentials resolve through SecretsManager per voice. The system prompt is back to being about reasoning, not formatting.

Delivery failures don’t crash the schedule timer. They surface as a schedule:delivery_failed event you can subscribe to. The job keeps running. The next morning it tries again.

A small thing, but it removes an entire layer of bespoke plumbing per agent.

2. Self-improving skills: the agent learns from itself

The runtime watches successful agent runs. Each run produces a trajectory — the sequence of tools called, in order, with hashed inputs. After the same shape of trajectory repeats N times (configurable; the shipped default is 5), Tutti asks the score’s LLM to summarise it into a skill candidate: a name, a description, the constituent tools, and an inner system prompt that does the whole sequence in one go.

The candidate is not a tool yet. It sits in the SkillStore waiting for review:

$ tutti-ai skills proposed
Pending skill candidates (1):

  • review_pr  (code-reviewer)
    Review a pull request: fetch the diff, read each changed file,
    and post a consolidated review comment.
    constituents: get_pull_request, get_file_contents, comment_on_issue
    evidence:     5 trajectories

$ tutti-ai skills review
> Review candidate review_pr? [a]pprove / [e]dit / [r]eject / [s]kip

Approve, and on the next run the agent has review_pr available as a single tool. The LLM picks it instead of orchestrating the three constituents itself. Internally, an inner-loop sub-agent runs against the approved system prompt and the same constituent tools — the outer agent sees one tool call where it used to see three.

Turning it on is a two-line score change:

import { InMemorySkillStore } from "@tuttiai/skills";

export default defineScore({
  // ...agents
  skills: {
    enabled: true,
    auto_propose_threshold: 5,
  },
});

export const skillStore = new InMemorySkillStore();

The runtime cost when skills are off is zero — the observer, proposer, and executor are only constructed when skills.enabled === true and you pass a store.

Why an operator gate? The proposer is good but not infallible. It can lump weakly-related flows into one skill, or propose a “skill” that’s actually a coincidence — three identical sequences that don’t generalise. Approving narrowly is the only way to keep the skill library trustworthy. Permissions on the synthesised skill are the union of constituents’ required_permissions, checked at run start. Rejected candidates are kept with the rejection reason, so the same shape doesn’t get re-proposed twice.

The point isn’t an agent that improves autonomously. The point is an agent that proposes its own improvements, and a workflow that lets a human accept them in five seconds instead of designing them in five hours.

3. Serverless deploy with a hibernation precheck

Most agents are bursty. They should pay only when they run. Tutti’s deploy bundler now ships a Modal target:

tutti-ai deploy --target modal

The generated modal_app.py runs the Tutti Node server inside node:20-bookworm-slim, installs tutti-ai@latest, declares each manifest.secrets entry as a Modal secret, and exposes the server via @modal.web_server(port=3000, startup_timeout=120). Scale-to-zero by default; idle costs nothing.

The hard part isn’t the bundler — it’s the contract. A serverless function isn’t a server. Between invocations, everything in memory evaporates. The naive port of a working Tutti score to Modal silently breaks: an InMemorySessionStore loses every session per cold start; an in-process node-cron schedule stops firing because the process is gone; a voice’s client cache opens fresh connections on every wake.

We could have papered over this with footnotes scattered through the docs. Nobody reads scattered footnotes. Instead, there’s a precheck:

$ tutti-ai deploy verify-hibernate
✗ score holds long-lived state that won't survive a cold start:
  - InMemorySessionStore is in use (sessions evaporate per cold start)
  - schedule on agent "reporter" runs on a node-cron timer (won't fire while hibernated)

Switch to a persistent SessionStore (e.g. PostgresSessionStore) and host
schedules in Modal's @modal.Function(schedule=…) instead of the in-process scheduler.

A static analysis pass that exits non-zero on a finding. Run it in CI before tutti-ai deploy --target modal. It’s not exhaustive — we can’t catch every closure that someone wedged into a voice constructor — but it covers the common shapes that bite first. Daytona ships alongside as a different point on the same curve: a dev sandbox that hibernates after idle and auto-resumes on traffic.


The shape that connects them

Look at the three score-level surfaces:

Each one is a declarative statement about operational intent. Each one is implemented by code the user didn’t have to write: a delivery dispatcher that knows how to talk to five platforms, a trajectory observer + proposer + reviewer + executor loop, a Modal bundler with a hibernation precheck. The agent’s code — its system prompt, its tools, its reasoning — stays simple.

This is the design principle Tutti has been converging on for the last few releases: the agent should not have to know about its own operations. Permissions belong to the runtime, not the agent. Scheduling belongs to the runtime, not the agent. Delivery belongs to the runtime. Hibernation rules belong to the runtime. The agent gets to be about reasoning. Everything else is a config block.

When you stand back, all three primitives in this post are the same idea pressed against three different problems. The framework is doing more so the agent does less.


Try it

npm i -g tutti-ai
tutti-ai init my-agent

Each primitive has a runnable example and a doc page:

What’s on the bench

Each of these has a follow-up worth flagging:

If any of those would unblock a real workload you’re trying to ship — or if your team is hitting a fifth “after-the-demo” problem we haven’t addressed yet — open a GitHub issue. The hardest signal to act on is silence.

Older post
Inbox: every channel, one agent
9 min · Engineering

Start conducting.

One install. Your first agent running in 60 seconds. No signup. No telemetry.