TL;DR: AWS developer interviews in 2026 lean hardest on Lambda: cold starts, concurrency, event sources and failure handling, then on DynamoDB, SQS, IAM roles and VPC networking. Candidates should know that Lambda now bills the INIT phase, that SnapStart covers Python and .NET as well as Java, and that Lambda Managed Instances and durable functions launched in late 2025.
Lambda durable functions, launched on December 2, 2025, let one function checkpoint its progress and pause for up to a year. That covers work that used to need Step Functions.
Two days earlier, Lambda Managed Instances began running functions on EC2 capacity in your own account. Both change the standard answer to "when would you not use Lambda?"
- 1Each Lambda function can add 1,000 execution environments every 10 seconds, up to the account's concurrency limit (1,000 by default).
- 2Current Lambda runtimes include Node.js 24, Python 3.14, Java 25, .NET 10 and Ruby 4.0; Node.js 20 was deprecated on April 30, 2026.
- 3DAX passes strongly consistent reads straight to DynamoDB and does not cache them.
- 4The X-Ray SDKs and daemon entered maintenance mode on February 25, 2026; new tracing work uses OpenTelemetry.
AWS Lambda
1. What is a Lambda cold start, and how do you reduce it?
A cold start is the extra time Lambda takes when no warm execution environment is free. Lambda must create one, download the code, start the runtime and run your code outside the handler. Warm requests skip all of that.
The fixes, in the order to try them:
- Trim the init code. Fewer and lighter dependencies, and lazy imports for rarely used paths.
- SnapStart. Lambda snapshots the initialized environment when you publish a version and resumes from it. It supports Java 11+, Python 3.12+ and .NET 8+, per the SnapStart docs.
- Provisioned concurrency. Pre-initialized environments that respond "in double-digit milliseconds", billed while they are on.

Layers do not make cold starts faster. The code in a layer still loads into the same environment. Measure before fixing: the REPORT log line shows Init Duration for every cold start.
2. What is the difference between reserved and provisioned concurrency?
Reserved concurrency caps and guarantees how many environments one function may use; provisioned concurrency keeps some of them warm. They solve different problems.
Reserved concurrency takes part of the account limit for one function. No other function can use it, and the function cannot go above it. That protects a database behind the function, and it stops one busy function from starving the rest.
It costs nothing.
- A cap and a guarantee for one function
- No extra charge
- Does not keep anything warm
- Pre-initialized environments on a version or alias
- Billed while configured
- Removes cold starts up to that count
Setting reserved concurrency to 0 is a quick way to stop a function from running at all.
3. Which Lambda limits come up in design reviews?
Timeout, memory, payload size and scaling rate. The Lambda quotas page lists them:
- Timeout: 15 minutes (90 minutes for async and event source mapping calls on Managed Instances).
- Memory: 128 MB to 10,240 MB. CPU grows with memory, and 1,769 MB equals one vCPU.
- Payload: 6 MB for synchronous requests and responses, 1 MB for asynchronous events, 200 MB for a streamed response.
/tmp: 512 MB to 10,240 MB.- Package: 50 MB zipped upload, 250 MB unzipped with layers, 10 GB for a container image.
- Scaling: 1,000 new environments every 10 seconds per function.
A payload over 6 MB is the usual reason to hand the client a presigned S3 URL instead of passing data through Lambda.
4. What are the three ways Lambda gets invoked, and how do retries differ?
Synchronous, asynchronous and event source mapping, and each handles errors differently. With a synchronous call (API Gateway, a function URL, the SDK's RequestResponse), the caller gets the error and decides whether to retry.
With asynchronous invocation (S3 events, SNS, EventBridge), Lambda queues the event and retries failed runs twice by default. It can then send the event to an on-failure destination.
With event source mapping (SQS, Kinesis, DynamoDB Streams, Kafka), Lambda polls the source and invokes in batches.
Retries follow the source: for SQS the message returns after the visibility timeout, and for streams the batch is retried until it succeeds or expires.
5. How do you stop one bad SQS message from failing a whole batch?
Turn on partial batch responses. Set ReportBatchItemFailures in the event source mapping's FunctionResponseTypes, and return the IDs of the messages that failed. Only those return to the queue.
def handler(event, context):
failures = []
for record in event["Records"]:
try:
process(record)
except Exception:
failures.append({"itemIdentifier": record["messageId"]})
return {"batchItemFailures": failures}
Without it, one failure makes Lambda retry the whole batch, so successful messages run again. The SQS error handling docs note that the Powertools for AWS Lambda batch utility handles this logic for you.
6. What are Lambda layers for, and what are their limits?
A layer is a ZIP archive of libraries, a custom runtime or shared files that several functions can use. Its contents appear under /opt in the execution environment.
A function can use up to 5 layers, and the function plus its layers must fit in 250 MB unzipped. Layers help share code and keep uploads small.
They also have costs: versions must be pinned, local testing is harder, and they do not work with container image functions, which bundle dependencies in the image instead.
7. Why must Lambda handlers be idempotent?
Because Lambda can run the same event more than once. Asynchronous invocations retry on error, SQS standard queues can deliver a message twice, and a timeout after the work finished looks like a failure to the caller.
The usual pattern stores a key for each processed event, such as an order ID or a hash of the payload, in DynamoDB with a conditional write. A repeat event finds the key and returns the saved result instead of charging a card twice.
Powertools for AWS Lambda ships an idempotency utility built on this pattern.
8. What changes when a Lambda function runs inside a VPC?
The function can reach private resources such as RDS, ElastiCache or internal services, but it loses default internet access. Outbound internet calls then need a NAT gateway in a public subnet.
Calls to AWS services can go through VPC endpoints instead.
A common interview trap is the old claim that VPC functions have much worse cold starts. Lambda now attaches functions through shared network interfaces created when the function is configured, not on each cold start.
The real VPC questions today are NAT cost, subnet IP space and security group rules.
Containers and EC2
9. When would you choose ECS on EC2, Fargate or EKS?
Pick by how much of the host and the orchestrator you want to own. ECS is AWS's own container scheduler, and EKS runs Kubernetes. Fargate is a way to run tasks or Pods from either one without managing servers.

- ECS on Fargate: the shortest path for a team that wants containers and no cluster work.
- ECS or EKS on EC2: for GPUs, custom AMIs, daemons on every node, or steady load where Savings Plans on your own fleet cost less.
- EKS: when the company already runs Kubernetes, or needs its tools, such as Helm, operators and service meshes. EKS Auto Mode now manages the nodes as well.
10. What are Spot Instances, and when do you use them?
Spot Instances are spare EC2 capacity at up to 90% off On-Demand prices, and AWS can take them back. EC2 sends an interruption notice two minutes before it stops or ends the instance, per the Spot interruption docs.
Use them for work that can stop and resume: batch jobs, CI runners, stateless web tiers behind a load balancer, and big data clusters.
Spread requests across several instance types and Availability Zones so one pool running dry does not stop everything. An Auto Scaling group with a mix of On-Demand and Spot keeps a floor of capacity that cannot be taken away.
Databases and Storage
11. How does Aurora differ from standard RDS?
Aurora splits compute from storage. Its storage is one shared cluster volume with copies of the data across three Availability Zones, per the Aurora storage docs. Standard RDS attaches block storage to each instance.
Because replicas read the same volume, Aurora read replicas lag less and a failover promotes one quickly without copying data. Aurora also adds features RDS lacks, such as Global Database, fast cloning and Aurora Serverless v2.
It costs more, so a small single-AZ app may be better on plain RDS.
12. What is Aurora Serverless v2, and can it scale to zero?
Aurora Serverless v2 adjusts database capacity in small steps, measured in Aurora capacity units (ACUs), as load changes. You set a minimum and maximum, and you pay for the capacity in use.
It can now scale to zero: set the minimum to 0 ACUs and an idle instance pauses automatically, with no instance charge while paused, per the auto-pause docs.
The first connection after a pause waits for it to resume, so auto-pause suits dev and test databases more than production APIs.
13. How do you design a DynamoDB partition key?
Choose a key with many distinct values that requests hit evenly, so load spreads across partitions. A user ID or order ID usually works. A status field or today's date does not, because most traffic lands on one value.
Design the keys from the access patterns: list every query the app makes, then pick partition and sort keys, plus global secondary indexes, that serve each one without a scan.
For a key that must take heavy writes, add a suffix (write sharding) and read across the shards.
14. What is DAX, and what does it do with strongly consistent reads?
DAX is an in-memory, write-through cache for DynamoDB with the same API, so reads can come back in microseconds. It caches eventually consistent reads in an item cache and a separate query cache.
Strongly consistent reads are the catch. Per the DAX consistency docs, DAX passes them to DynamoDB and "does not cache the results". Writes made directly to DynamoDB, bypassing DAX, leave stale items in the cache until the TTL expires.
The query cache is not updated by item writes at all.
15. Explain the cache-aside pattern with ElastiCache.
The app reads from the cache first and, on a miss, reads the database and writes the result into the cache with a TTL. Only data that someone requests gets cached.
def get_user(user_id):
cached = cache.get(f"user:{user_id}")
if cached:
return json.loads(cached)
user = db.fetch_user(user_id)
cache.set(f"user:{user_id}", json.dumps(user), ex=300)
return user
The weak spots are stale data after a write and a stampede when a hot key expires. Deleting the key on write and adding jitter to TTLs handle both. ElastiCache now supports Valkey alongside Redis OSS and Memcached.
16. How do S3 conditional writes help concurrent writers?
A conditional write succeeds only if a precondition on the object holds, so two clients cannot silently overwrite each other. With If-None-Match: *, a PutObject fails with 412 Precondition Failed if the key already exists.
With If-Match and an ETag, it fails if someone changed the object since you read it.
S3 added conditional writes on August 20, 2024. Bucket policies can now require them. Before that, teams needed a DynamoDB lock table to get the same safety.
Networking and APIs
17. What is the difference between a security group and a network ACL?
Security groups are stateful allow-lists on network interfaces; network ACLs are stateless rule lists on subnets. With a security group, return traffic for an allowed request is allowed automatically, and rules can only allow.
Network ACL rules can allow or deny, are checked in number order, and need explicit rules for return traffic, including ephemeral ports. In practice, security groups do most of the work.
A security group can also allow traffic from another security group, which is cleaner than listing IP ranges.
18. What are VPC endpoints, and which kind should you use?
VPC endpoints let resources in private subnets reach AWS services without an internet gateway or NAT gateway. There are two main kinds. Gateway endpoints, for S3 and DynamoDB, are route table entries and cost nothing.
Interface endpoints (PrivateLink) place a network interface in your subnet for most other services and are billed per hour and per GB.
A NAT gateway charges for every GB it processes, so moving heavy S3 traffic to a gateway endpoint is one of the quickest savings on an AWS bill. Endpoint policies can also limit which buckets or actions are reachable.
19. When do you choose API Gateway REST APIs over HTTP APIs?
Choose REST APIs when you need their extra features, and HTTP APIs when you do not, because HTTP APIs cost less. Per the API Gateway comparison, only REST APIs offer:
- API keys, usage plans and per-client throttling
- AWS WAF, private endpoints and resource policies
- Caching, request validation and request body transformation
- X-Ray tracing, canary deployments and response streaming
HTTP APIs have a built-in JWT authorizer and automatic deployments, and they suit simple Lambda or HTTP backends.
20. What is a Lambda authorizer?
A Lambda authorizer is a function API Gateway calls before the backend to decide whether a request is allowed. It receives the token or request details.
For REST APIs it returns an IAM policy, and HTTP APIs can also accept a simple true or false response.
Results can be cached by token for a set time, which cuts cost and latency. Use one for custom schemes, such as an API key checked against your own database.
For standard OIDC tokens, an HTTP API's JWT authorizer or a Cognito authorizer needs no code.
Messaging and Events
21. Compare SQS, SNS and EventBridge.
SQS is a queue, SNS is a push-based topic, and EventBridge is an event bus with content-based routing. In SQS, consumers poll, and each message is handled by one consumer.
In SNS, every subscriber gets a copy of each message, which is the classic fan-out.
EventBridge routes events from AWS services, SaaS partners and your own apps to targets. Rules match on any field in the event, and it adds schema discovery, archive and replay.
A common design publishes domain events to EventBridge and gives each consumer its own SQS queue as a buffer.
22. What is the difference between SQS standard and FIFO queues?
Standard queues give at-least-once delivery with best-effort ordering; FIFO queues keep order within a message group and avoid duplicates.
Per the FIFO docs, SQS drops a retried send with the same deduplication ID within a 5-minute interval.
The deduplication ID comes from you or from a SHA-256 hash of the body with content-based deduplication. Ordering applies per MessageGroupId, so use many groups, such as one per customer, to process in parallel.
Standard queues suit most work when consumers are idempotent.
23. Explain the SQS visibility timeout and dead-letter queues.
A received message stays hidden for the visibility timeout. If the consumer deletes it in time, it is gone; if not, it becomes visible again for another try.
Set the timeout longer than the worst-case processing time. With Lambda, AWS advises at least six times the function timeout.
A redrive policy moves a message to a dead-letter queue after a set number of receives, so a message that always fails stops blocking work. Alarm on the DLQ depth, and use redrive to send fixed messages back.
24. How does Kinesis Data Streams differ from SQS?
Kinesis is an ordered, replayable log; SQS is a queue that deletes messages once handled. Many consumers can read the same Kinesis records on their own, each tracking its position, and ordering holds within a shard.
Records stay for 24 hours by default, extendable up to 365 days, per the retention docs. Use Kinesis for clickstreams, telemetry and change feeds that several systems read. Use SQS for task distribution, where each job should run once.
IAM and Security
25. How do IAM users, roles and policies relate?
Policies are JSON documents of permissions, and users and roles are identities that policies attach to. A user has long-term credentials.
A role has none; trusted principals assume it and get temporary credentials from STS through sts:AssumeRole.
Workloads should always use roles: an execution role for Lambda, a task role for ECS, an instance profile for EC2. People should sign in through IAM Identity Center rather than as IAM users with access keys.
Cross-account access works the same way: a role in account B trusts account A, and A's principals assume it.
26. When do you use Secrets Manager rather than Parameter Store?
Use Secrets Manager for credentials that need rotation, and Parameter Store for configuration and simple secrets.
Secrets Manager can rotate database passwords on a schedule with built-in or custom Lambda rotation, and it can replicate secrets across Regions. Each secret has a monthly cost.
Parameter Store's standard tier is free and stores SecureString values encrypted with KMS. It suits feature flags, endpoints and API keys that rarely change.
Both are read at runtime through IAM, so nothing secret sits in code or the CI config.
27. How does KMS use envelope encryption?
KMS generates a data key, you encrypt the data with it locally, and KMS encrypts the data key under a KMS key that never leaves the service. You store the encrypted data key next to the data.
To decrypt, you send the encrypted data key to KMS, get the plaintext key back, and decrypt locally. This keeps large data out of KMS API calls, which have size and rate limits.
Each use of the KMS key is checked by IAM and the key policy and logged in CloudTrail. S3, EBS and RDS encryption all work this way.
Operations and Cost
28. CloudFormation or CDK?
Both end in CloudFormation. CDK is a library that generates CloudFormation templates from TypeScript, Python, Java, C# or Go. CloudFormation is the engine that creates and updates the resources.
CDK adds loops, types, tests and higher-level constructs, such as a queue wired to a function with the right permissions. Plain templates are easier to review line by line and need no build step.
A strong candidate knows that cdk diff and change sets show what will change before a deploy. Stack drift and resource replacement are the real production risks with either tool.
29. How do you trace a request across AWS services?
Instrument the services with OpenTelemetry and send traces to X-Ray or another backend. X-Ray builds a service map and shows where time goes across API Gateway, Lambda, DynamoDB and other calls.
The instrumentation layer changed recently: the X-Ray migration guide says the X-Ray SDKs and daemon entered maintenance mode on February 25, 2026, with releases limited to security fixes.
Pair traces with structured logs that carry the trace ID, so one ID finds every log line for a request.
30. How would you cut the AWS bill for a running system?
Start with Cost Explorer to find the biggest line items, then fix those first. The usual wins:
- Right-size instances and databases using real usage, not launch-day guesses.
- Cover steady load with Savings Plans, and move interruptible work to Spot.
- Replace NAT gateway traffic to S3 and DynamoDB with gateway endpoints.
- Add S3 lifecycle rules or Intelligent-Tiering for old data, and delete unattached volumes and old snapshots.
- Switch suitable workloads to Graviton instances or
arm64Lambda.
Tag resources by team and service first, or nobody can say who owns the spend. The Well-Architected Framework's cost optimization pillar is one of its six.
What Changed Recently
31. Why does the Lambda INIT phase show up on bills now?
Since August 1, 2025, Lambda bills the INIT phase for every configuration. Before that, on-demand ZIP functions on managed runtimes got it free.
The AWS Compute Blog says INIT time now counts toward billed duration at the standard on-demand rate.
For most apps the cost is small, because cold starts are a small share of invocations. It matters for functions with heavy init code and spiky traffic. It also ended an old trick of doing expensive work at init time to get it for free.
32. How does SnapStart work, and what can break with it?
Lambda runs your init code when you publish a version, snapshots the memory and disk state, and resumes new environments from that snapshot. It reached Python and .NET on November 18, 2024, after starting with Java.
Anything unique created during init is shared by every environment restored from the snapshot: random seeds, UUIDs, temporary credentials. Network connections opened at init may be stale after a restore.
SnapStart does not work with provisioned concurrency, EFS or /tmp above 512 MB. It runs only on published versions, not $LATEST. Python and .NET pay caching and restore charges; Java does not.
33. What are Lambda Managed Instances?
They run Lambda functions on EC2 instances in your account that Lambda manages, launched on November 30, 2025. You pick instance types through a capacity provider. Lambda handles patching, routing and scaling.
Per the Managed Instances docs, one environment handles many requests at once, so handler code must be thread-safe. Scaling follows CPU use "without cold starts", but the function does not scale to zero.
Pricing is the EC2 price plus a 15% management fee, and Savings Plans apply to the EC2 part. It fits steady, high-volume traffic.
34. What are Lambda durable functions, and when would you use Step Functions instead?
Durable functions let one Lambda function run a multi-step workflow for up to a year, with checkpoints and waits that cost nothing while suspended. They were announced on December 2, 2025.
Per the durable functions docs, after a pause the code replays from the start and skips steps that already have checkpoints. Code outside a step must therefore be deterministic. SDKs exist for JavaScript, TypeScript, Python and Java.
Step Functions still fits orchestration across many AWS services, with a visual designer and direct integrations to over 220 services.
35. Which Lambda runtimes should a team plan to leave?
Anything on Amazon Linux 2, plus Node.js 20 and older. The Lambda runtimes page lists Amazon Linux 2 end of life as June 30, 2026.
Python 3.10, Python 3.11, Java on AL2 and provided.al2 get only limited patches until their own deprecation dates.
Node.js 20 was deprecated on April 30, 2026, and Node.js 18 on September 1, 2025. New functions should use Node.js 24, Python 3.13 or 3.14, Java 21 or 25, or .NET 10.
After deprecation, AWS blocks creating and then updating functions on the old runtime, so an unplanned upgrade can block an urgent fix.
Signs of a Strong Answer
- They measure
Init Durationbefore choosing a cold start fix, and know SnapStart's uniqueness caveat. - They make every event handler idempotent without being asked, and can say where duplicates come from.
- They set the SQS visibility timeout and partial batch responses correctly with Lambda.
- They design DynamoDB keys from access patterns and know when DAX will not help.
- They reach for roles and temporary credentials everywhere, never long-lived access keys in code.
- They can explain a line on the bill, such as NAT gateway data processing, and how to remove it.
Hiring AWS Developers
Engineers who can run serverless systems on AWS in production are in demand, and many strong ones work remotely from Asia.
Second Talent matches companies with pre-vetted DevOps and cloud engineers and back-end developers, screened with questions like these.
Tell us the stack and we send a shortlist within 24 hours. Start hiring, or see our Azure and GCP interview guides.






