Skip to content

Top 30 Event-Driven Architecture Interview Questions and Answers [2026]

September 29, 2026 15 min read
Event-Driven Architecture interview guide by Second Talent
TL;DR: Event-driven architecture interviews test whether a candidate can keep data correct when messages arrive late, twice or out of order. Expect questions on sagas, the outbox pattern, idempotent consumers, schema evolution, and when to pick a queue or a log.

RabbitMQ 4.0, tagged in September 2024, removed classic queue mirroring. Kafka 4.2 then made queue-style share groups production-ready in February 2026.

The line between "queue" and "log" brokers is blurrier than it was, so interviewers now ask why a design needs each one. The patterns that decide most interviews, such as sagas, outboxes and idempotent consumers, have not changed.

Key takeaways
  1. 1A saga does not roll back. It runs compensating steps, and other readers can see the half-done state in between.
  2. 2RabbitMQ 4.0 quorum queues ship with a default redelivery limit of 20, after which a message is dead-lettered or dropped.
  3. 3AsyncAPI 3 replaced the confusing publish and subscribe keywords with send and receive. The current version is 3.1.0.
  4. 4EventStoreDB is now KurrentDB, renamed in version 25.0 in March 2025.

Core Concepts

1. What is the difference between an event, a command and a message?

An event states a fact that already happened, such as OrderPlaced. A command asks one recipient to do something, such as PlaceOrder, and it may be refused. A message is the envelope that carries either one.

The difference shows up in ownership and naming. Events are named in the past tense and owned by the producer. Any number of consumers may react, and the producer does not know who they are.

Commands are named as orders and are owned by the receiver's API. A design that publishes "events" like SendWelcomeEmail is sending commands through a topic. That couples the two services while hiding it.

2. What are the main styles of events?

Martin Fowler's "What do you mean by Event-Driven?" names four patterns that often get mixed up:

  • Event notification: a small event, often only an ID. Consumers call back for details, which adds load and coupling to the source.
  • Event-carried state transfer: the event holds the data consumers need, so they can keep a local copy and skip the callback.
  • Event sourcing: the events are the system of record, and state is rebuilt from them.
  • CQRS: separate models for writes and reads, often kept in sync by events.

A strong candidate asks which one a team means before designing anything. The trade-offs differ a lot.

3. What are the benefits of an event-driven architecture?

Loose coupling in time and in knowledge. The producer does not wait for consumers, and it does not need to know they exist.

  • A new consumer can subscribe to existing events without a change to the producer.
  • Producers and consumers scale on their own. A slow consumer builds a backlog instead of slowing the producer.
  • If a consumer is down, the broker holds its events until it comes back.
  • With a log broker, a new service can replay history to build its own view.

4. What are the main challenges of event-driven architecture?

Consistency, visibility and change. Each benefit above has a cost.

  • Eventual consistency: other services see a change later, and user interfaces must handle that gap.
  • Duplicates and ordering: most brokers deliver at least once, and order holds only within a partition or queue.
  • Hidden flow: no single place shows the whole business process, so debugging needs tracing and good naming.
  • Contract drift: an event schema is a public API. Changing it can break consumers you did not know about.
  • Error handling: there is no caller to return an error to, so failures need retries, dead letter queues and alerts.

5. What is eventual consistency, and how do you design for it?

Eventual consistency means that, once updates stop, all copies of the data converge to the same value. Until then, parts of the system may read stale data.

You design for it in the product as much as in code. Show a "processing" state rather than pretending the work is done. Read your own writes from the service that owns them.

Make sure business rules that must never break sit inside one service's transaction.

A payment ledger should not depend on two services agreeing through events. An answer that treats the delay as a small technical detail usually misses these product effects.

6. When would you choose orchestration over choreography?

Choose orchestration when the process has many steps, branches or deadlines, and someone needs to see where each case stands. Choose choreography for short flows where each service reacts to an event and nothing needs a central view.

Choreography
  • Services react to events
  • No central state; the flow lives in many codebases
  • Suits short flows between bounded contexts
Orchestration
  • A coordinator sends commands
  • One place holds each case's state and timeouts
  • Suits long processes with branches and deadlines

Workflow engines such as Temporal, Camunda or AWS Step Functions play the coordinator role. A common rule is choreography between bounded contexts, and orchestration inside a long business process.

Patterns

7. How does the saga pattern keep data consistent across services?

A saga splits one business transaction into local transactions, one per service. Each step commits and then triggers the next. If a step fails, the saga runs compensating transactions that undo the earlier steps.

Swimlane of a choreographed order saga: the order service marks the order pending, payment charges the card, inventory reports out of stock, then payment issues a refund and the order is cancelled.

Compensation is not a rollback. The card was charged and is later refunded. In between, other services can see the order as pending and the payment as taken. The microservices.io saga page lists this lack of isolation as the main drawback.

Good answers name the defences: semantic locks such as a PENDING status, steps ordered so the hardest to undo runs last, and compensations that are idempotent.

8. What problem does the transactional outbox solve?

It removes the dual write. Without it, a service commits to its database and then publishes to the broker. If it crashes between the two, the data and the event disagree.

With the outbox pattern, the service writes the business row and an outbox row in the same local transaction. A relay then publishes the outbox rows.

The relay can poll the table, or read the database's change log with a CDC tool such as Debezium.

BEGIN;
INSERT INTO orders (id, status) VALUES ('o-42', 'PENDING');
INSERT INTO outbox (id, aggregate_id, type, payload)
  VALUES ('e-901', 'o-42', 'OrderPlaced', '{"orderId":"o-42"}');
COMMIT;

The relay can publish the same row twice after a crash. So the outbox gives at-least-once delivery, and consumers still need to be idempotent.

9. What is event sourcing, and when is it worth it?

Event sourcing stores every change to an entity as an event, and rebuilds current state by replaying them. The event log is the source of truth, not a table of current rows.

It pays off where the history matters to the business: ledgers, audits, bookings, or any domain where "how did we get here?" is a real question. It costs more elsewhere. Events are hard to change once written.

Queries need separate read models. Long streams need snapshots to load quickly. Martin Fowler's Event Sourcing article is the usual starting point. Event sourcing is also not the same as publishing events.

Most event-driven systems do the second without the first.

10. How does CQRS fit with events?

CQRS uses one model to handle commands and separate models to serve queries. Events are how the query models learn about changes.

A command updates the write model and emits an event. Projection handlers consume the event and update read models shaped for each screen or API, such as a search index or a summary table.

The cost is lag: a user may not see their change right away. Teams handle it by returning the new state from the command, or by waiting for the projection to catch up.

Fowler's CQRS note warns that it adds real complexity and suits only some parts of a system.

11. How do you make a consumer idempotent?

Make processing the same event twice have the same effect as processing it once. There are three common ways.

  • Natural idempotency: write absolute values, such as "set status to SHIPPED", rather than "add 1".
  • Processed-ID table (an inbox): store each event ID in the same transaction as the business change, with a unique constraint. A duplicate fails the insert and is skipped.
  • Version checks: apply an update only if the event's version is newer than the stored one.

The key detail is atomicity. If the ID is stored in Redis while the business change goes to Postgres, a crash between the two brings the duplicate back.

Brokers and Delivery

12. How do log-based brokers differ from message queues?

A queue hands each message to one consumer and removes it once it is acknowledged. A log appends messages and keeps them for a retention period. Each consumer group tracks its own position.

Decision flowchart: use a log if several services need the same events or you need replay, use a queue if tasks need per-message acks, routing or priority, otherwise pick what the team runs well.

That difference drives everything else. Logs such as Kafka and Pulsar support many independent consumers, replay and high throughput. Ordering holds per partition.

Queues such as RabbitMQ and SQS offer per-message acknowledgement and delays, and RabbitMQ adds routing and priorities. They suit task distribution. The line has blurred.

RabbitMQ has streams, which are logs, and Kafka 4.2 made share groups production-ready, which behave like queues.

13. What do at-most-once, at-least-once and exactly-once mean in practice?

At-most-once may lose messages but never repeats them. At-least-once never loses them but may repeat them. Exactly-once means each message's effect happens once.

At-least-once is the normal choice, paired with idempotent consumers. Exactly-once exists only inside certain boundaries. Kafka transactions, for example, make a read-process-write flow atomic within Kafka.

As soon as a consumer calls an external API or writes to another database, the guarantee ends, and idempotency has to cover that step. Treat "exactly-once" in a design review as a question: exactly-once across which systems?

14. How do you keep events in order?

Route all events for one entity through one ordered channel. In Kafka that means using the entity ID as the key, so they share a partition. In SQS FIFO it means sharing a message group ID.

Global order across all entities is rarely needed and does not scale. Within an entity, retries can still reorder events if a consumer processes messages in parallel.

Defences include a sequence number or version on each event, so consumers can drop stale ones. Handlers that give the same result in any order also help.

15. What is a dead letter queue, and how do you run one well?

A dead letter queue holds messages that failed processing after a set number of attempts. It keeps a single bad message from blocking everything behind it.

A DLQ is only useful if someone looks at it. Good practice:

  • Retry transient errors with backoff first. Send permanent errors, such as a schema mismatch, straight to the DLQ.
  • Keep the original payload and add the error, source and attempt count as headers.
  • Alert on DLQ depth and age, not only on its existence.
  • Build a replay tool, and fix the cause before replaying.
A default worth knowing. Since RabbitMQ 4.0, quorum queues redeliver a message at most 20 times by default. After that it goes to the dead letter target, or it is dropped if none is set. So a quorum queue with no dead letter policy can lose poison messages quietly.

16. How are competing consumers different from publish-subscribe?

With competing consumers, many instances of one service share a queue, and each message goes to one of them. That is how you scale out one kind of work. With publish-subscribe, every subscribing service gets its own copy of each event.

Kafka does both through consumer groups. Instances in one group compete for partitions, and different groups each receive every event. The number of partitions caps how many instances in a group can work at once.

Share groups remove that cap for queue-style work.

17. How do you handle a consumer that cannot keep up?

First measure it: consumer lag, in messages and in time, is the signal. Then either speed up the consumer or add capacity.

Options include scaling out instances (up to the partition count on a log broker), batching writes to the database, and moving slow calls off the main path. Some consumers can also shed work, such as skipping old analytics events.

A broker buffers the backlog, but only until its retention or disk limit. So a lagging consumer needs an alert well before events start to expire.

Schemas and Contracts

18. How do you evolve an event schema without breaking consumers?

Make only compatible changes, and let a schema registry enforce them. Confluent's schema evolution docs define the modes most teams use:

  • Backward: consumers on the new schema can read data written with the old one. Upgrade consumers first.
  • Forward: consumers on the old schema can read data written with the new one. Upgrade producers first.
  • Full: both at once. Usually this means only adding or removing optional fields with defaults.

Renaming a field or changing its type is a breaking change under every mode. For those, publish a new event type or a new topic, and run both for a while.

19. How do you handle old event versions in an event-sourced system?

Upcast them when they are read. An upcaster turns version 1 of an event into version 2 in memory, so the domain code only handles the latest shape.

The stored events stay as they were, because the log is the record of what happened. Rewriting the store ("copy and replace") is possible but risky. Teams keep it for rare cases, such as removing personal data.

Weak answers suggest editing stored events in place.

20. What is CloudEvents?

CloudEvents is a CNCF specification for common event metadata. It defines required attributes such as id, source, specversion and type, plus bindings for HTTP, Kafka, AMQP and other transports.

It standardizes the envelope, not the payload. A gateway, broker or tracing tool can route and log events from any producer without knowing their business schema. The current core spec is version 1.0.2. The main site is cloudevents.io.

21. What is Event Storming, and what does it produce?

Event Storming is a workshop where developers and domain experts map a business process as a timeline of domain events on sticky notes. Commands, actors, policies and external systems are then added around the events.

The output is a shared picture of the process and its hot spots. It also shows where the language changes, which hints at bounded context borders and so at service boundaries.

Alberto Brandolini created the method, and eventstorming.com is its home page.

Testing and Operations

22. How do you trace a request across asynchronous events?

Put the trace context in the message headers when publishing, and restore it when consuming. The W3C Trace Context traceparent header is the usual format, and OpenTelemetry messaging instrumentation handles it for common clients.

Async flows add a twist. The OpenTelemetry messaging conventions use span links, not parent and child spans, as the default way to tie a consumer's work to the producer's context.

That fits batches, where one receive span links to many producers. Candidates should also mention correlation IDs in the event body for business-level lookups, which outlive trace retention.

23. What should you monitor in an event-driven system?

Consumer lag comes first, measured as the age of the oldest unprocessed event. It shows user-visible delay directly.

  • Publish and consume rates per topic or queue.
  • End-to-end latency from event time to processing time.
  • Error and retry rates per consumer, and DLQ depth and age.
  • Broker health: disk, replication, and partitions or queues without a leader.

24. How do you test an event-driven service?

Test each service at its boundary: feed it events and check the events and state it produces. Unit tests cover handlers as plain functions.

Integration tests run a real broker in a container, for example with Testcontainers, to catch serialization and configuration bugs that mocks hide.

Contract tests check that producer and consumer agree on the schema, either through a schema registry's compatibility check or a tool like Pact. End-to-end tests across many services are slow and flaky.

Keep a few for the most important flows.

25. How do you replay events safely?

Replay into consumers that are idempotent and that do not repeat outside effects. Rebuilding a read model is safe. Resending customer emails is not.

Common tactics: build the new projection in a fresh table and switch reads when it catches up. Mark replayed messages with a header so handlers can skip side effects. On Kafka, use a new consumer group or reset offsets for one group.

Check the retention first, because you can only replay what the broker still holds.

What Changed Recently

Jun 17, 2024
CloudEvents SQL 1.0 (CESQL) released
Sep 2024
RabbitMQ 4.0: classic queue mirroring removed, AMQP 1.0 core
Mar 25, 2025
KurrentDB 25.0: EventStoreDB renamed
Nov 18, 2025
Axon Framework 5.0: async-native APIs
Jan 31, 2026
AsyncAPI 3.1.0 released
Feb 17, 2026
Kafka 4.2: share groups production-ready

26. What did RabbitMQ 4.0 change for event-driven designs?

It removed classic queue mirroring. Classic queues now have one replica, so replicated data must use quorum queues or streams. The RabbitMQ 4.0 release notes list the other changes that matter:

  • AMQP 1.0 is a core protocol that is always on.
  • Quorum queues have a default redelivery limit of 20.
  • Quorum queues now support message priorities.
  • Khepri, the new metadata store that replaces Mnesia, is fully supported.

A cluster that relied on mirroring policies still runs after the upgrade, but without replication. That is the trap an interviewer wants the candidate to spot.

27. How do Kafka share groups change the queue versus log choice?

They let Kafka act as a work queue. Consumers in a share group read the same partitions at once and acknowledge each record on its own. The Kafka 4.2 announcement declared them production-ready.

A consumer can accept, release or reject each record, and Kafka counts delivery attempts. Parallelism is no longer capped by the partition count. The trade-off is ordering, which share groups do not keep.

So a team can run both event streams and task queues on one platform. Whether it should depends on whether it needs RabbitMQ's exchange-based routing, which Kafka does not offer.

28. What changed in AsyncAPI 3?

Operations were split from channels, and publish and subscribe were replaced with send and receive. In version 2, it was never clear whose point of view publish described.

Per the 3.0 release notes, channels now describe only the addresses and messages, so they can be reused across documents. Version 3 also added request and reply.

Version 3.1.0, released on January 31, 2026, added ROS 2 bindings, per the 3.1.0 release notes.

asyncapi: 3.1.0
channels:
  userSignedUp:
    address: user.signedup
    messages:
      userSignedUp:
        $ref: '#/components/messages/UserSignedUp'
operations:
  sendUserSignedUp:
    action: send
    channel:
      $ref: '#/channels/userSignedUp'

29. What changed for event-sourcing databases and frameworks?

Two of the main tools changed. EventStoreDB was renamed KurrentDB in version 25.0, released on March 25, 2025. Axon Framework 5.0 followed on November 18, 2025.

Per the KurrentDB 25.0 release notes, the rename covers the executable, configuration prefixes, HTTP headers and metrics. The release also added archiving of old chunks to S3.

The Axon Framework 5.0 notes describe async-native command handling, event stores and event processors. They also add event tags and append conditions, used for Dynamic Consistency Boundaries.

Axon's notes say migration tooling from version 4 was planned for 5.1.

30. What does Kafka 4.0 removing ZooKeeper mean for an event platform team?

One less system to run, and a forced upgrade path. Kafka 4.0, released on March 18, 2025, runs only in KRaft mode, where Kafka's own controllers store cluster metadata.

Per the 4.0 announcement, a ZooKeeper-based cluster must first migrate to KRaft on the 3.9 bridge release. The same release made the new consumer rebalance protocol generally available.

For consumers that opt in, it removes the stop-the-world pause during rebalances, which used to stall event consumers during deploys.

Signs of a Strong Answer

  • They separate events from commands by who owns the contract, not only by naming.
  • They say a saga compensates rather than rolls back, and name how they handle the visible in-between state.
  • They reach for an outbox the moment a service writes to a database and publishes an event.
  • They make consumers idempotent with a stored ID in the same transaction as the change.
  • They pick the ordering key from the business entity, and say what happens to order on retry.
  • They treat event schemas as public APIs, with a compatibility mode and a plan for breaking changes.

Hiring Event-Driven Engineers

Event-driven systems fail in subtle ways, so the right hire has run one in production.

Second Talent matches companies with pre-vetted back-end engineers from Asia who have built on Kafka, RabbitMQ and cloud queues, screened with questions like these.

Tell us the stack and we send a shortlist within 24 hours. Start hiring, or see our Kafka and domain-driven design interview guides.

Hiring developers in Southeast Asia?

Get Cost Guide

How would you like to talk?

WhatsApp us Prefer texting at your own pace? Just hit us up on WhatsApp. We promise no spam and a hassle-free experience.

Loading available times…