This is the story of a Telegram bot that started as one function and became a layered system. No big design upfront. Just small refactoring steps.


Day 1. From scratch

Before. 8804b50

The first step for me was trying to figure out how a Telegram bot works, so I didn’t choose any framework, I just used a pure fetch method to basically have a ping.

// This is the entry point into Cloudflare Worker
export default {
  async fetch(req: Request, env: Env): Promise<Response> {
    if (req.method !== 'POST') {
      return new Response('OK')
    }

    const update = await req.json()
    const chatId = update?.message?.chat?.id

    await fetch(`https://api.telegram.org/bot${env.BOT_TOKEN}/sendMessage`, {
      method: 'POST',
      headers: { 'Content-Type': 'application/json' },
      body: JSON.stringify({ chat_id: chatId, text: 'Hello, world =)' }),
    })

    return new Response('OK')
  },
}

After this I tried adding more security. But I realized that testing that in a manual way is extremely inefficient. So I needed a way to do this automatically with tests.

Problem. HTTP validation, authorization, Telegram API call. All in one block. To test the security check you have to run the whole handler. To test the reply you have to mock fetch globally. Nothing is isolated.


Step 1. Extract a function.

After. 3679644

function isRequestValid(req: Request, env: Env): boolean {
  if (req.method !== 'POST') return false
  const secret = req.headers.get('x-telegram-bot-api-secret-token')
  return secret === env.BOT_WEBHOOK_SECRET
}

What changed. The security check is now a pure function. It takes a Request, returns a boolean. No side effects.

Test.

it('should reject GET requests', () => {
  const noWebhookSecret = new Request('http://example.com')

  const validity = isRequestValid(getRequest, env)

  expect(validity).toBe(false)
})

it('should reject POST without webhook secret', () => {
  const noWebhookSecret = new Request('http://example.com', { method: 'POST' })

  const validity = isRequestValid(noWebhookSecret, env)

  expect(validity).toBe(false)
})

No mock needed. No bot. No database. Three lines per case.


Step 2. Extract processBot(). Inject the dependency.

Before. b53f5e1

async function processBot(update, env: Env) {
  const chatId = update?.message?.chat?.id
  const userId = update?.message?.from?.id
  if (chatId && userId === Number(env.ADMIN_USER_ID)) {
    await fetch(`https://api.telegram.org/bot${env.BOT_TOKEN}/sendMessage`, {
      method: 'POST',
      headers: { 'Content-Type': 'application/json' },
      body: JSON.stringify({ chat_id: chatId, text: 'Hello, world =)' }),
    })
  }
}

Problem. To test whether the bot responds to the right users, you have to intercept a real HTTP call to api.telegram.org. The dependency — how to send a message — is hardcoded inside the function.

After. f32c342

async function processBot(
  update: unknown,
  env: Env,
  sendMessage: (chatId: number, text: string) => Promise<void>
) {
  const chatId = update?.message?.chat?.id
  const userId = update?.message?.from?.id
  if (chatId && userId === Number(env.ADMIN_USER_ID)) {
    await sendMessage(chatId, 'Hello, world =)')
  }
}

What changed. sendMessage is now a parameter. The caller decides what sending means. In production it calls Telegram. In tests it calls a mock.

Before the injection (b53f5e1) — tests went through the full HTTP handler:

global.fetch = vi.fn()  // intercept ALL fetch calls globally

it('responds with OK for non-admin user', async () => {
  const request = new IncomingRequest('http://example.com', {
    method: 'POST',
    body: JSON.stringify({
      message: { chat: { id: 12345 }, from: { id: 99999 } },
    }),
    headers: {
      'Content-Type': 'application/json',
      'x-telegram-bot-api-secret-token': 'test-webhook-secret',
    },
  })

  await worker.fetch(request, env)

  expect(mockFetch).not.toHaveBeenCalled()
})

Problem. Three things are mixed in one test. The transport — constructing an HTTP request with the right headers and JSON body. The environment — the webhook secret must be correct or the test fails before reaching the user logic. The user behavior — the actual thing being tested, whether a non-admin user gets a response.

To test one condition about a user ID, you had to get the HTTP layer right first.

After the injection (654fe39) — tests call processBot directly:

const mockSendMessage = vi.fn()

it('should not respond to non-admin user', async () => {
  await processBot(
    { message: { chat: { id: 12345 }, from: { id: 99999 } } },
    testEnv,
    mockSendMessage
  )

  expect(mockSendMessage).not.toHaveBeenCalled()
})

it('should respond to admin user', async () => {
  await processBot(
    { message: { chat: { id: 12345 }, from: { id: 54321 } } },
    testEnv,
    mockSendMessage
  )

  expect(mockSendMessage).toHaveBeenCalledWith(12345, 'Hello, world =)', testEnv)
})

No HTTP. No JSON. No global mock. Each test touches exactly one concern.


Step 3. Adopt a bot framework. Keep the seam.

Before. f32c342 — manual JSON parsing, manual routing, raw fetch to send replies.

Problem. Routing and serialization are boilerplate. They grow with every new command. The business logic is buried in parsing code.

After. d408606

bot.on('message:text', async (ctx) => {
  await processMessage(ctx.message?.text, async (text, other) => {
    await ctx.reply(text, other)  // grammy-specific, stays here
  })
})

What changed. grammy handles routing and serialization. processMessage receives plain text and a reply callback. It still doesn’t know what grammy is.

The seam is preserved: reply is injected. The framework lives at the boundary. The handler stays testable.


Step 4. Add a port. Name it explicitly.

Before. 6a7b70d

async function processMessage(
  text: string,
  reply: (text: string) => Promise<void>
) {
  if (text.startsWith('/new_person')) {
    // no way to save to DB — the feature can't be implemented
    await reply('✅ Added person: ...')
  }
}

Problem. The function needs to write to a database but has no way to do it without taking a direct dependency on D1. Passing D1Database would tie the business logic to the infrastructure.

After. 4fd179c (interface) + 95d429e (implementation)

interface PeopleClient {
  addPerson(firstName: string, lastName: string): Promise<void>
}

async function processMessage(
  text: string,
  reply: (text: string) => Promise<void>,
  peopleClient: PeopleClient,
) { ... }

In production, the real implementation:

class LivePeopleClient implements PeopleClient {
  constructor(private readonly _d1: D1Database) {}

  async addPerson(firstName: string, lastName: string) {
    await this._d1.prepare('INSERT INTO ...').bind(...).run()
  }
}

Wired up in index.ts — the only place that knows about D1:

const peopleClient = new LivePeopleClient(env.DB)

bot.on('message:text', async (ctx) => {
  await processMessage(ctx.message?.text, async (text, other) => {
    await ctx.reply(text, other)
  }, peopleClient)
})

In tests, three kinds of fake — each for a different purpose.

A stub for tests where PeopleClient is not the subject:

const stubPeopleClient = {} as PeopleClient

A resolving mock to test the happy path — did the right thing get called?

const resolvingPeopleService: PeopleService = {
  addPerson: vi.fn(() => Promise.resolve()),
}

A rejecting mock to test the error path — does the handler recover correctly?

const rejectingPeopleService: PeopleService = {
  addPerson: vi.fn(() => Promise.reject('some error from DB')),
}

Tests for both paths:

it('should create person with first and last name', async () => {
  await processMessage('/new_person John Smith', mockReply, resolvingPeopleService)

  expect(resolvingPeopleService.addPerson).toHaveBeenCalledWith('John', 'Smith')
  expect(mockReply).toHaveBeenCalledWith('✅ Added person: John Smith')
})

it('should show error when person creation fails', async () => {
  await processMessage('/new_person John Smith', mockReply, rejectingPeopleService)

  expect(mockReply).toHaveBeenCalledWith('❌ Error creating person: some error from DB')
})

The stub is used when the client doesn’t matter for what’s being tested. The resolving mock tests the good path. The rejecting mock tests the error path. Both matter. You are not done testing a feature until you have tested what happens when it fails.

What changed. The function now depends on an interface, not on D1. The test passes whatever it wants through that interface. D1 is not involved.


Step 5. Name infrastructure explicitly.

Before. 598997b

class MilongaRepository { ... }   // is this the interface or the D1 class?
class EventStore { ... }
class UnitOfWork { ... }

Problem. The name MilongaRepository could mean the interface (what the application needs) or the implementation (how D1 does it). They’re the same class. You can’t swap the implementation. You can’t write a fake.

After. 856f5ee (MilongaRepository) + 4047487 (D1EventStore)

// Port: what the application needs
interface MilongaRepository {
  getOpenedMilongaId(): Promise<string | null>
  getAvailableGuests(milongaId: string): Promise<Person[]>
  countUnpaidGuests(milongaId: string): Promise<number>
}

// Infrastructure: how it's done with D1
class D1MilongaRepository implements MilongaRepository {
  constructor(private readonly _d1: D1Database) {}

  async getOpenedMilongaId(): Promise<string | null> {
    return (
      await this._d1
        .prepare('SELECT milonga_id FROM milonga_read WHERE closed_at IS NULL LIMIT 1')
        .first<{ milonga_id: string }>()
    )?.milonga_id ?? null
  }
}

What changed. The D1 prefix makes the distinction visible in the name. Interface = port. Class with D1 prefix = infrastructure. You can now write a FakeMilongaRepository for tests and swap it in.

But before the split, tests looked like this. d29bfb8

it('should not check-in same guest twice', async () => {
  await db
    .prepare('INSERT INTO people_read (person_id, first_name, surname, created_at) VALUES (?, ?, ?, ?)')
    .bind('p1', 'John', 'Doe', Date.now())
    .run()

  await client.openMilonga('admin')
  await client.checkInGuest('admin', 'p1')

  await expect(client.checkInGuest('admin', 'p1')).rejects.toThrow(
    'D1_ERROR: UNIQUE constraint failed: events.idempotency_key: SQLITE_CONSTRAINT'
  )
})

Three problems. The test needs a real database running. You have to open a milonga and check in a guest just to reach the line you actually want to test. And the error is a raw D1 constraint — the business rule is expressed as an infrastructure artifact. You only discover this error after deploying and running the app live.

After the split, the same rule. 22bdb68

it('should not check-in same guest twice', async () => {
  await client.openMilonga('admin')
  await client.checkInGuest('admin', 'p1')

  await expect(client.checkInGuest('admin', 'p1')).rejects.toThrow('Guest already checked in')
})

No database setup. Two lines of arrange. The error is a sentence a human wrote. The test reads like a story.

The length of the test is the smell. A long test means something is mixed that shouldn’t be. When you need a database to check a business rule, the rule is in the wrong place.

As the business rules moved into this layer, the names changed too. LiveMilongaClient became DefaultMilongaService. LivePeopleClient became DefaultPeopleService. 52ab8bc + 22a07e9

// before: milongaClient — sounds like a DB wrapper
const milongaClient = new DefaultMilongaService(...)
actionContext.milongaClient

// after: milongaService — it enforces business rules, not just fetches data
const milongaService = new DefaultMilongaService(...)
actionContext.milongaService

A client fetches data. A service enforces invariants. The rename touched 36 files — every handler, every test. But it was a mechanical rename, not a rewrite. The name had been wrong for a while. The code was already right.


Step 6. Remove data validation from handlers.

Before. 59081bf — handlers parse their own arguments and validate them:

export async function handleGuestCheckIn(ctx: CallbackContext) {
  const { text, fromUser, reply, milongaClient } = ctx

  const commandData = text.split(':')
  if (commandData.length !== 2) {
    return  // silent failure
  }

  const personId = commandData[1]

  try {
    await milongaClient.checkInGuest(fromUser, personId)
    // ...
  } catch (e) {
    if (e instanceof Error) {
      if (e.message.includes('UNIQUE constraint failed') || e.message.includes('idempotency_key')) {
        await reply('ℹ️ This guest is already checked in.')
      }
    } else {
      await reply(`❌ Error checking in guest: ${e}`)
    }
  }
}

Problem. Three concerns mixed in one function: argument parsing, argument validation, and business logic. Every handler repeated the same parsing pattern. And the D1 error leaked up — the handler had to know what a UNIQUE constraint failed error meant.

After. 037aeb4 — handler receives parsed args, knows nothing about encoding:

export async function handleGuestCheckIn(ctx: ActionContext) {
  const { args, fromUser, reply, milongaService } = ctx
  const [personId] = args

  try {
    await milongaService.checkInGuest(fromUser, personId)
    await reply('✅ Guest checked in', { reply_markup: ... })
  } catch (e) {
    await reply(`❌ Error checking in guest: ${e}`)
  }
}

What changed. The handler receives args: string[] — already split, already validated. It doesn’t know the callback was encoded as check_in:personId. That’s the bot layer’s job.

The parsing and argument validation moved into createCallback() — one place, once:

// src/bot/callback.ts
export function createCallback(
  process: (ctx: ActionContext) => Promise<void>,
  argCount: number,
  prefix: string
): Callback {
  return {
    prefix,
    process: async (ctx) => {
      const { args } = ctx
      if (args.length !== argCount) return  // wrong shape → ignore

      args.forEach((arg, index) => {
        if (arg.length === 0) {
          throw new Error(`Invalid argument #${index} for ${prefix}. Expected non-empty argument.`)
        }
      })

      await process(ctx)
    },
    encode(...args: string[]) {
      if (args.length !== argCount) {
        throw new Error(`Invalid argument count for ${prefix}. Expected ${argCount}, got ${args.length}`)
      }
      return args.length === 0 ? prefix : `${prefix}:${args.join(':')}`
    },
  }
}

processCommand and processCallback became thin dispatchers — they look up the handler and call it. Nothing else:

// src/bot/process-command.ts
export async function processCommand(
  commandRegistry: CommandsRegistry,
  trigger: string,
  ctx: ActionContext
) {
  const command = commandRegistry[trigger]
  if (command) {
    await command.process(ctx)
  }
}

// src/bot/process-callback.ts
export async function processCallback(
  callbackRegistry: CallbackRegistry,
  prefix: string,
  ctx: ActionContext
) {
  const callback = callbackRegistry[prefix]
  if (callback) {
    await callback.process(ctx)
  }
}

Tests. processCommand and processCallback are tested independently of any real command or callback. A mock handler is injected:

// test/interaction/process-command.test.ts
const mockProcess = vi.fn()
const testCommand = createCommand({ trigger: '/match', process: mockProcess })
const commands = buildCommandRegistry([testCommand])

it('should not process unknown command', async () => {
  await processCommand(commands, '/not_match', stubContext)
  expect(mockProcess).not.toHaveBeenCalled()
})

it('should process matching command', async () => {
  await processCommand(commands, '/match', stubContext)
  expect(mockProcess).toHaveBeenCalled()
})

Argument validation in createCallback is tested by itself:

// test/interaction/process-callback.test.ts
const oneArgCallback = createCallback(mockProcess, 1, 'one_arg')
const testHandlers = buildCallbackRegistry([oneArgCallback])

it('should not call handler with missing argument', async () => {
  await processCallback(testHandlers, 'one_arg', mockContext([]))
  expect(mockProcess).not.toHaveBeenCalled()
})

it('should not call handler with empty argument', async () => {
  await expect(
    processCallback(testHandlers, 'one_arg', mockContext(['']))
  ).rejects.toThrow('Invalid argument #0 for one_arg. Expected non-empty argument.')
})

it('should call handler with correct argument', async () => {
  await processCallback(testHandlers, 'one_arg', mockContext(['arg1']))
  expect(mockProcess).toHaveBeenCalledWith(expect.objectContaining({ args: ['arg1'] }))
})

The encode side is tested too — the same abstraction that parses also encodes:

it('should encode single argument', () => {
  const { encode } = createCallback(vi.fn(), 1, 'test')
  expect(encode('arg1')).toEqual('test:arg1')
})

it('should throw on argument count mismatch', () => {
  const { encode } = createCallback(vi.fn(), 1, 'test')
  expect(() => encode('arg1', 'extra')).toThrow('Invalid argument count for test. Expected 1, got 2')
})

Step 7. Separate the folders.

Before. 645808d — all files at src/: index.ts, milongaClient.ts, peopleClient.ts, eventStore.ts — flat, growing.

Problem. The folder structure doesn’t communicate what anything does. To understand the system you have to read every file.

After. Current state.

src/
  infrastructure/d1/    ← D1EventStore, D1MilongaRepository, D1UnitOfWork, ...
  ports/                ← EventStore, MilongaRepository, UnitOfWork, ...
  application/          ← DefaultMilongaService, DefaultPeopleService
  interaction/          ← handleMilongaOpen, showMilongaStatus, ...
  bot/                  ← commands, callbacks, process-command.ts
  index.ts              ← entry point + security

What changed. The folder is the layer. Where a file lives tells you what it knows and what it’s allowed to depend on. Infrastructure knows D1. Application knows ports. Interaction knows services. Bot knows grammy.

This is not a framework. It’s naming and moving files, one refactor at a time.


What you can test at each layer.

Each layer has a different kind of test, because each layer has different dependencies.

Security layer. Pure functions. No mocks.

it('should reject POST without webhook secret', () => {
  expect(isRequestValid(new Request('http://x.com', { method: 'POST' }), env)).toBe(false)
})

Bot layer. Just verify the wiring is correct.

it('should match the command', () => {
  expect(milongaOpenCommand.trigger).toEqual('/milonga_open')
})

Interaction layer. Mock the service interface. Check what reply was called with.

it('should open a milonga and show guest list', async () => {
  await handleMilongaOpen({
    args: [],
    fromUser: 'mock-user-id',
    reply: mockReply,
    milongaService: resolvingMilongaService,
    peopleService: stubPeopleService,
  })

  expect(resolvingMilongaService.openMilonga).toHaveBeenCalledWith('mock-user-id')
  expect(mockReply).toHaveBeenCalledWith('Milonga opened')
})

Application layer. Fake repositories. No D1.

class FakeMilongaRepository implements MilongaRepository {
  _milongaOpened = false

  getOpenedMilongaId() {
    return Promise.resolve(this._milongaOpened ? 'fake-id' : null)
  }
  // ...
}

Or plug in real D1 for integration tests. The service doesn’t know the difference.


The pattern.

There is no single most important step here. All of them matter equally. It’s like physical growth — there is no most important centimeter. You just grow.

What drives each step is pain. Something becomes hard to change, hard to test, hard to read. That pain is a signal. You listen to it, you make one small move, and the system becomes a little easier to work with.

Kent Beck said it well: make the change easy, then make the easy change. The first part is the hard part. That’s what this whole sequence is about.

Every step was small. Extract a function. Add a parameter. Name a class. Move a file. No step required understanding the full system. No step required a rewrite.

The tests got simpler with each step, not harder. When a test is long, something is mixed that shouldn’t be. When a test needs a database to check a business rule, the rule is in the wrong place. When a mock becomes complex, the abstraction is wrong. The tests tell you.

Each layer ended up testing exactly what it owns — with exactly the kind of fake that layer deserves. Security: plain inputs, no mocks. Bot: trigger matching. Interaction: mocked services, checked replies. Application: fake repositories, no DB. Infrastructure: real D1, snapshot.

The architecture was not designed upfront. It emerged from the refactoring. The refactoring was guided by the tests. The tests were guided by pain.

That’s evolvability. Not a property you add at the end. A habit you build one step at a time.