TL;DR: Senior Lambda interviews in 2026 test what happens under load and under failure: execution environment phases and their limits, concurrency math, async retry tuning, and how one bad record stalls a Kinesis shard. Candidates should also know the 2025 and 2026 changes: SQS provisioned mode, tenant isolation mode, the Node.js 24 handler rules and the 200 MB streaming limit.
On August 3, 2026, AWS raised the ceiling for SQS provisioned mode to 10,000 event pollers, enough for 100,000 concurrent invocations from one queue mapping. Standard mode still stops at 1,250.
The questions below sit one level under the basics in our AWS interview guide. They ask a candidate to reason about limits, failure paths and cost.
- 1With default stream settings, one bad Kinesis record can block its shard for up to one week.
- 2Throttled asynchronous events go back on Lambda's internal queue and are retried for up to 6 hours by default.
- 3SQS messages that fail an event filter are deleted from the queue, not left for another consumer.
- 4Recursive loop detection stops a chain after about 16 invocations, but cannot see loops that pass through DynamoDB.
Execution Environment
1. Walk through the execution environment lifecycle. What are the time limits in each phase?
There are three phases: Init, Invoke and Shutdown, and each has its own limit. Per the execution environment docs, Init starts the extensions, bootstraps the runtime and runs your static code, and it is limited to 10 seconds.
If it does not finish in time, Lambda retries Init at the first invocation, this time under the function timeout.

The 10-second limit does not apply with provisioned concurrency, SnapStart or Lambda Managed Instances. There, init may run for 130 seconds or the function timeout, whichever is higher.
The function timeout limits Invoke, and it covers the handler and every extension together.
Shutdown gets 0 ms with no extensions, 500 ms with an internal extension and 2,000 ms with one or more external extensions. After that, Lambda sends SIGKILL.
A strong candidate adds that a failed Init writes an INIT_REPORT line with a status such as timeout or Extension.Crash. That line is the first place to look when a function never reaches its handler.
2. What is a suppressed init, and how do you spot one?
It is the re-initialization Lambda runs after an invoke fails. A crash, timeout or out-of-memory error during Invoke makes Lambda reset the environment, and the next request on it runs Init again before the handler.
Lambda does not log a separate INIT phase for it. Instead, the Duration in the REPORT line includes the extra init time, so it no longer matches the gap between the START and REPORT timestamps.
The docs' example shows a 2.5-second gap next to a reported 3,022.91 ms. To see suppressed inits directly, subscribe an extension to the Telemetry API. It emits INIT_START, INIT_RUNTIME_DONE and INIT_REPORT events with phase=invoke.
In practice, a function that times out often enough pays for init over and over, and its duration graphs look worse than its handler code explains.
3. What survives between invocations, and what must you never rely on?
Anything created outside the handler survives while the environment stays warm: SDK clients, database connections, caches and files in /tmp. That reuse is a free optimization, not a guarantee.
/tmpoutlives a reset. The docs state that a reset after an invoke failure does not clear/tmpbefore the next Init. A half-written file from a crashed run can still be there.- Environments are recycled. Lambda ends execution environments every few hours for runtime updates and maintenance, even for functions invoked all the time.
- Init can run early. For on-demand functions, Lambda sometimes initializes environments ahead of requests, so there can be a gap between Init and the first invoke. AWS says not to depend on this.
Good answers check that a reused connection is still alive before using it, write temp files under a per-request name, and keep no state that only one environment knows about.
4. What belongs in init code, and what belongs in the handler?
Put work in init that every request needs and that stays valid for the life of the environment. That means SDK clients, connection pools, parsed config and compiled templates.
Put anything per-request, or anything that expires, in the handler.
import boto3
# Init: runs once per environment, counted in Init Duration
dynamodb = boto3.resource("dynamodb")
table = dynamodb.Table("orders")
def handler(event, context):
# Invoke: per-request work only
order_id = event["order_id"]
return table.get_item(Key={"pk": order_id}).get("Item")
Two senior details. First, init is billed and must finish in 10 seconds on demand. A slow secret fetch or a big model load belongs behind a lazy cache, not at import time.
Second, short-lived values, such as a session token read from the environment, should be resolved in the handler.
The Parameters and Secrets extension docs make the same point: resolve the token inside the handler so it stays fresh after a SnapStart restore.
Concurrency and Scaling
5. How do you estimate the concurrency a workload needs?
Multiply average requests per second by average duration in seconds. The concurrency docs give the formula and a worked case: 5,000 requests per second at 200 ms each needs a concurrency of 1,000.
Each in-flight request gets its own execution environment, so concurrency is also the number of environments.
That one example already uses up the default account limit of 1,000 concurrent executions per Region, shared by every function in the account.
The follow-up is what to do about it. Cut the duration with faster downstream calls or more memory. Ask for a quota increase. Or move part of the load to a batch path. Halving duration halves concurrency at the same traffic.
Latency work on a hot function is also capacity work.
6. How fast can one function scale during a spike?
Each function can add up to 1,000 execution environments every 10 seconds, or 10,000 requests per second every 10 seconds, per the scaling behavior docs. The rate applies per function, so two functions scale independently.
Unused scaling capacity does not build up. A function that was quiet for a minute still gets only 1,000 new environments in the next 10 seconds.
Requests that arrive faster than Lambda can scale, or after the function hits its concurrency limit, fail with a 429 throttling error.
For a known launch or a scheduled batch, provisioned concurrency initializes environments before the traffic arrives.
7. What happens to a throttled request under each invocation type?
It depends on who holds the event. With a synchronous call, the 429 goes straight back to the caller, and the client or API Gateway decides whether to retry.
With an asynchronous call, Lambda already has the event. Per the async error handling docs, throttles (429) and system errors (500-series) send the event back to the queue.
Lambda keeps retrying for up to 6 hours by default, with backoff that grows from 1 second to at most 5 minutes.
With an event source mapping on a stream, the Kinesis error handling docs cover throttling. A batch that cannot be invoked is retried until its records expire or pass MaximumRecordAgeInSeconds. The shard waits behind it the whole time.
8. How do SQS maximum concurrency and reserved concurrency interact?
They are separate settings, and setting them inconsistently causes throttling. Maximum concurrency caps how many function instances one SQS event source mapping can invoke, from 2 to 1,000, at no charge.
Reserved concurrency caps the function as a whole.
The SQS scaling docs say reserved concurrency must be greater than or equal to the total maximum concurrency of all SQS mappings on the function. If it is lower, the docs warn that Lambda might throttle the messages.
Maximum concurrency is the right tool when two queues feed one function and neither should starve the other. Reserved concurrency is the right tool when the database behind the function has a hard connection limit.
Async and Failures
9. How would you tune async invocation for a payment webhook?
Shorten the event age, decide the retry count on purpose, and always add an on-failure destination. The two knobs, per the async configuration docs, are MaximumEventAgeInSeconds (up to 6 hours) and MaximumRetryAttempts (0 to 2).
aws lambda put-function-event-invoke-config \
--function-name payment-webhook \
--maximum-event-age-in-seconds 3600 \
--maximum-retry-attempts 1 \
--destination-config '{"OnFailure":{"Destination":"arn:aws:sqs:us-east-1:123456789012:payment-failed"}}'
By default a function error is retried twice, one minute after the first attempt and two minutes after the second.
A payment confirmed six hours late can be worse than one that fails fast and lands in a queue for a human, so a lower event age often makes sense.
One trap: put-function-event-invoke-config overwrites the whole config, while update-function-event-invoke-config changes only the fields you pass.
10. On-failure destination or dead-letter queue: which, and why?
Use an on-failure destination for new work. It accepts more targets and records why the event failed. A DLQ only receives the original event.
- Standard SQS queue or SNS topic only
- Gets the event content, no response details
- Set at function level; versions share the $LATEST setting
- SQS, SNS, S3, another function or an EventBridge bus
- Invocation record with condition, invoke count and error
- Set per function, version or alias
The invocation record carries requestContext.condition (such as RetriesExhausted), approximateInvokeCount, the request payload and the error from the response.
FIFO queues and FIFO topics are not valid targets; a failed delivery shows up only in the DestinationDeliveryFailures metric. SNS has a 256 KB message limit, so a near-1 MB event plus metadata can fail to deliver.
For large payloads the docs point to SQS or S3.
A follow-up worth asking: what does setting reserved concurrency to 0 do to async events? Lambda sends new events straight to the DLQ or on-failure destination with no retries.
It is a kill switch that does not lose data, provided a destination exists.
11. How do you add idempotency with Powertools, and what happens when the first attempt times out?
Wrap the handler with the Powertools @idempotent decorator and a DynamoDB persistence layer. It hashes the chosen part of the event and stores the result.
A repeat within the expiry window, one hour by default, gets the stored result, per the Powertools idempotency docs.
from aws_lambda_powertools.utilities.idempotency import (
DynamoDBPersistenceLayer, IdempotencyConfig, idempotent,
)
persistence = DynamoDBPersistenceLayer(table_name="IdempotencyTable")
config = IdempotencyConfig(
event_key_jmespath="powertools_json(body).order_id",
expires_after_seconds=24 * 60 * 60,
)
@idempotent(config=config, persistence_store=persistence)
def handler(event, context):
return charge_card(event)
The senior part is concurrency and timeouts. While the first call is still running, a duplicate gets IdempotencyAlreadyInProgressError rather than a second charge.
If the first call times out, the lock would normally block retries until the record expires. The decorator handles this by storing the remaining invocation time.
For @idempotent_function on an inner function, you call config.register_lambda_context(context) yourself.
Picking the key matters as much as the library: hash the order ID, not the whole body, or a retry with a new timestamp header looks like a new request.
Event Sources
12. How does the SQS event source mapping scale in standard mode?
It starts small and adds pollers as the queue fills. Per the SQS scaling docs, Lambda begins with five batches on five concurrent invokes and adds up to 300 more concurrent invokes per minute while messages remain.
One mapping tops out at 1,250 concurrent invokes.
As traffic drops, it scales back to five, and it can go as low as 2 to save on SQS API calls. Setting maximum concurrency turns that optimization off.
The 1,250 ceiling also explains a common surprise: an account whose concurrency quota was raised to 5,000 still cannot drain one queue faster than 1,250 invokes in standard mode. Provisioned mode (question 25) is the answer to that.
13. How do FIFO queues limit Lambda concurrency?
Concurrency is capped at the number of active message group IDs or the maximum concurrency setting, whichever is lower. Six message groups with maximum concurrency of 10 gives at most six concurrent invocations.
A batch can hold messages from several groups, and order is kept within each group. If the function returns an error, Lambda retries the affected messages before it reads more from that group, so one failing message pauses its whole group.
Design follows from this: choose a group key with many values, such as customer ID, so that parallelism comes from the number of groups. Pair it with partial batch responses so that one failure does not replay the rest of the batch.
14. One bad record is stuck in a Kinesis stream. What happens, and how do you contain it?
Lambda keeps retrying the batch and the shard stops moving. Per the Kinesis error handling docs, with default settings a bad record can block its shard for up to one week.
The parameter reference shows why: MaximumRetryAttempts and MaximumRecordAgeInSeconds both default to -1, meaning no limit.

The containment toolkit:
ReportBatchItemFailures: return the sequence number of the first failed record. If you return several, Lambda checkpoints at the lowest one and retries from there.BisectBatchOnFunctionError: splits a failed batch in two to isolate the bad record. Splits do not use up the retry quota.- Retry and age limits: set
MaximumRetryAttempts(up to 10,000) andMaximumRecordAgeInSeconds(up to 604,800) to values your data can tolerate. - On-failure destination: SQS and SNS receive metadata only (stream ARN, shard ID, start and end sequence numbers). An S3 destination also receives the full invocation record.
The same parameters apply to DynamoDB Streams. Note the response shape: for streams, itemIdentifier is a sequence number, not an SQS message ID.
15. How does ParallelizationFactor work, and does it break ordering?
It lets Lambda process one shard with up to 10 concurrent batches (the default is 1) while keeping order per partition key. Per the Kinesis docs, a factor of 2 on 100 shards allows up to 200 concurrent invocations.
Use it when IteratorAge climbs and adding shards is slow or expensive. Records with the same partition key still arrive in order.
The catch is Kinesis aggregation: if the records inside an aggregated record have different partition keys, Lambda cannot guarantee order by partition key. Events that do not match the aggregated record's key are dropped and lost.
A candidate who uses the Kinesis Producer Library should raise this without prompting.
16. When does event filtering save money, and what happens to records that do not match?
Filtering saves money whenever a large share of records needs no work, because filtered records never invoke the function. A mapping gets up to five filters by default, raisable to 10 by quota request.
The filters are ORed, so a record that matches any one of them is delivered, per the event filtering docs.
Non-matching records are where teams get caught. For SQS, Lambda removes the message from the queue.
If another team expected to read those messages, they are gone, so a shared queue needs SNS or EventBridge fan-out in front of it instead of a filter.
For Kinesis and DynamoDB Streams, the iterator moves past the record, and it ages out under the stream's normal retention. Filtering is not supported for Amazon DocumentDB.
Security and Deploys
17. Execution role or resource-based policy: which one does each case need?
The execution role says what the function may do. The resource-based policy says who may invoke or manage the function.
Lambda assumes the execution role on every invocation, and AWS advises against calling sts:AssumeRole on it yourself in code.
- S3 or SNS triggers a function: a resource-based policy statement lets that service invoke it. Add
aws:SourceArnoraws:SourceAccountso that only your bucket can. - The function reads from SQS, Kinesis or DynamoDB Streams: the event source mapping polls with the execution role. That role needs permissions such as those in
AWSLambdaSQSQueueExecutionRole, pluskms:Decryptfor an encrypted queue.
Per the resource-based policy docs, a function's policy can be at most 20 KB. Statements added one at a time with AddPermission support only aws:SourceArn, aws:SourceAccount and aws:PrincipalOrgID as conditions.
18. How do you lock down a function URL?
Use AuthType: AWS_IAM unless the endpoint is meant to be public, and scope the policy to URL calls. With AWS_IAM, Lambda authorizes the caller against its identity policy and the function's resource-based policy.
With NONE, Lambda skips authentication, but the resource-based policy still has to grant public access.
A recent change catches older templates: per the function URL auth docs, new function URLs created from October 2025 need both lambda:InvokeFunctionUrl and lambda:InvokeFunction.
Add the lambda:InvokedViaFunctionUrl condition key to the InvokeFunction grant, or the principal can also call the function through other paths. For a NONE URL, the auth itself happens in the handler (a signed header or a token check).
IAM Access Analyzer flags functions that grant public or cross-account access, at no charge.
19. How should a function read secrets without calling Secrets Manager on every request?
Use the AWS Parameters and Secrets Lambda extension. It runs a local HTTP cache on port 2773 and fetches from Parameter Store or Secrets Manager only when the cached value expires.
import json, urllib.request, boto3
def handler(event, context):
token = boto3.Session().get_credentials().get_frozen_credentials().token
req = urllib.request.Request(
"http://localhost:2773/systemsmanager/parameters/get"
"?name=%2Faws%2Freference%2Fsecretsmanager%2Fprod-db"
)
req.add_header("X-Aws-Parameters-Secrets-Token", token)
db_config = json.loads(urllib.request.urlopen(req).read())
...
The details to know, from the extension docs: every request needs the X-Aws-Parameters-Secrets-Token header set to the session token. The TTLs (SSM_PARAMETER_STORE_TTL, SECRETS_MANAGER_TTL) default to and max out at 300 seconds.
The cache holds up to 1,000 items. Each environment keeps its own cache. The extension does not notice a rotated value until the TTL passes.
A rotation plan must allow for up to five minutes of the old secret, or keep the old credential valid for that window.
20. How do you roll out a new version safely, and who controls runtime patches?
Shift traffic gradually through an alias, and choose the runtime update mode on purpose. A weighted alias splits traffic between two published versions.
Both must share the same execution role and dead-letter queue, and the alias cannot point to $LATEST.
CodeDeploy automates the shift. Its predefined Lambda configs include LambdaCanary10Percent5Minutes (10% first, the rest five minutes later), LambdaLinear10PercentEvery1Minute and LambdaAllAtOnce.
With provisioned concurrency, raise it during the shift to avoid spillover onto cold environments.
Runtime patches are a separate track. Per the runtime management docs, the modes are Auto (the default), Function update and Manual. Auto uses a two-phase rollout: first to functions as they are created or updated, then to the rest.
Manual pins a runtime version ARN and is meant for rolling back a bad patch, not for long-term pinning, because a pinned function stops receiving security fixes.
Cost and Operations
21. How do you choose memory size and CPU architecture?
Measure, because memory also sets CPU. Lambda allocates CPU in proportion to memory, from 128 MB to 10,240 MB, and 1,769 MB equals one vCPU, per the memory docs.
A CPU-bound function can cost the same or less at a higher setting, when the run gets shorter by more than the price per millisecond goes up.
The open source AWS Lambda Power Tuning tool runs the function at several memory sizes through Step Functions and plots cost against speed. Compute Optimizer gives memory recommendations, but only for x86_64 functions.
For architecture, the architecture docs say arm64 (Graviton2) can give "significantly better price and performance" than x86_64.
The blockers are practical: every native dependency needs an arm64 build, and the package must be rebuilt for arm64. Run Power Tuning on both architectures before switching.
22. When does Fargate cost less than Lambda?
Fargate wins when the load is steady and high enough to keep a container fleet busy. The way to show it is the concurrency formula from question 5, not a rule of thumb.
Lambda bills per request plus GB-seconds of duration, and each in-flight request takes a whole environment. A Fargate task bills per second for vCPU and memory from image pull until it stops, with a one-minute minimum, busy or idle.
One task can serve many requests at once.
The working method in an interview:
- Compute average and peak concurrency from requests per second times duration, hour by hour, not as a daily average.
- Price Lambda at that concurrency and memory, from the Lambda pricing page.
- Price the number of tasks needed to serve the peak with headroom, plus the load balancer, from the Fargate pricing page.
- Add the work each option creates: container images, patching, autoscaling policies, idle capacity overnight.
Spiky or idle-heavy traffic favours Lambda, because Fargate pays for headroom it rarely uses. Flat, high traffic where requests spend most of their time waiting on I/O favours containers, which can share one process across many requests.
Lambda Managed Instances sits between the two; the AWS guide covers it.
23. How do you test and debug Lambda code beyond unit tests?
Test the handler logic locally, then test against real AWS services as early as possible, because IAM, event shapes and timeouts are where the bugs are. The AWS SAM CLI covers each step:
sam local invokeruns the function in a local Docker container with a sample event.sam sync --watchpushes code changes to a dev stack in the cloud as you save. AWS recommends it for development only andsam deployor CI/CD for production.sam remote invokecalls the deployed function from the terminal.
For bugs that only appear in the cloud, the AWS Toolkit for VS Code supports Lambda remote debugging. It adds a debugging layer, raises the timeout to 900 seconds and connects your local debugger over AWS IoT Secure Tunneling.
All changes are reverted afterwards. It supports Python, Node.js and Java, but not Managed Instances or container image functions. Use it on a dev copy of the function, never on the production alias.
24. How do you trace a request through Lambda now that the X-Ray SDKs are in maintenance mode?
Instrument with OpenTelemetry, usually through the AWS Distro for OpenTelemetry (ADOT) Lambda layer, and send traces to X-Ray or another backend.
The X-Ray migration guide states that the X-Ray SDKs and daemon entered maintenance mode on February 25, 2026, with security fixes only. AWS recommends OpenTelemetry for new instrumentation.
The ADOT Lambda docs describe layers that auto-instrument the function and turn on CloudWatch Application Signals by default. Set OTEL_AWS_APPLICATION_SIGNALS_ENABLED=false to keep plain OpenTelemetry tracing without it.
A senior candidate also pairs traces with JSON logs that carry the trace ID. X-Ray trace headers still matter inside Lambda, too: recursive loop detection (question 30) relies on them.
What Changed Recently
One more date for runtime questions: Rust support became generally available on November 14, 2025, backed by AWS Support and the Lambda SLA after years as "experimental".
Rust still deploys on an OS-only runtime, because it compiles to native code.
25. What is SQS provisioned mode, and when is it worth paying for?
It gives an SQS event source mapping a minimum and maximum number of dedicated event pollers, kept ready instead of scaled up on demand.
It launched on November 14, 2025, scaling 3x faster than standard mode, at up to 1,000 concurrent invokes per minute. On August 3, 2026 the maximum went from 2,000 to 10,000 pollers, for up to 100,000 concurrent invocations.
The minimum can be set from 2 to 200 pollers, and each poller handles up to 1 MB/s, 10 concurrent invokes or 10 SQS polling calls per second. Billing is per Event Poller Unit, on top of the SQS API calls the pollers make.
Provisioned mode cannot be combined with maximum concurrency; the maximum poller count becomes the concurrency control. For the lowest latency, the event source mapping docs say to set MaximumBatchingWindowInSeconds to 0.
It is worth paying for when bursts must clear in seconds, or when one queue needs more than standard mode's 1,250 invokes. Amazon MSK and self-managed Kafka mappings have their own provisioned mode.
26. What does tenant isolation mode guarantee, and what does it not support?
It guarantees that an execution environment serves only one tenant ID, so one tenant's code or data never shares an environment with another's. It was announced on November 19, 2025, for SaaS platforms that run tenant code or hold tenant data, which used to need one function per tenant.
The constraints, per the tenant isolation docs:
- It is set at creation (
TenantIsolationMode: PER_TENANT) and cannot be turned on for an existing function. - Every invoke must pass a tenant ID (
--tenant-id, or theX-Amz-Tenant-Idheader); without one, the invoke fails. The handler reads it fromcontext.tenantId. - It does not work with function URLs, provisioned concurrency or SnapStart. API Gateway REST APIs can map the header; HTTP APIs cannot.
- All tenants share one execution role, so the function's own code must still scope data access by tenant.
- There is a limit of 2,500 tenant-isolated environments per 1,000 concurrency, and an extra charge each time Lambda creates a tenant environment.
Event source mappings cannot pass the header. A June 2026 Compute Blog post shows the workaround: a small routing function reads the tenant ID from each message and invokes the isolated function with it.
27. What breaks when a function moves to the Node.js 24 runtime?
Callback-style handlers. The Node.js 24 launch post (November 25, 2025) says the new runtime interface client no longer supports the three-argument (event, context, callback) signature.
It also removes context.callbackWaitsForEmptyEventLoop, context.succeed, context.fail and context.done. Node.js 22 and earlier keep them.
// Node.js 22 and earlier only
export const handler = (event, context, callback) => {
doWork(event).then((r) => callback(null, r)).catch(callback);
};
// Node.js 24
export const handler = async (event: OrderEvent): Promise<Result> => {
return await doWork(event);
};
The quieter change concerns unresolved promises. Lambda no longer waits for them once the handler returns, and response streaming now behaves the same way once the stream ends.
Fire-and-forget work, such as an un-awaited analytics call, may never finish. A migration review should grep for callback, context.done and un-awaited promises, not only bump the runtime string.
28. What are the response streaming limits today?
A streamed response can be up to 200 MB, against 6 MB for a buffered one. The limit rose from 20 MB on July 31, 2025. The streaming docs add the details that change designs:
- The first 6 MB streams with no bandwidth cap; after that, the rate is capped at 2 MBps. A 200 MB response therefore takes well over a minute.
- A dropped client connection does not stop the function, and you pay for the full duration, so long timeouts on streaming functions cost money.
- Managed runtimes support streaming only on Node.js. Python and others need a custom runtime or the Lambda Web Adapter.
- Function URLs do not stream inside a VPC; there, use
InvokeWithResponseStreamthrough a Lambda interface endpoint.
API Gateway REST APIs gained response streaming on November 19, 2025, with request timeouts up to 15 minutes instead of the 29-second default. Streaming reached all commercial Regions on April 7, 2026.
Bytes streamed beyond the first 6 MB are billed as an extra charge.
29. How did Lambda logging costs change, and which controls cut them further?
Since May 1, 2025, Lambda logs in CloudWatch are billed as vended logs at volume tiers. In US East (N. Virginia), the rate starts at $0.50 per GB and falls to $0.05 per GB.
Lambda can also send logs to S3 or Data Firehose, priced from $0.25 down to $0.05 per GB. The Compute Blog's example: 60 TB a month drops from $30,000 to $12,500, 58% less.
Controls on top of the pricing:
- JSON format first. Per the log-level docs, log-level filtering requires JSON log format, and the default for managed runtimes is still plain text.
- Separate levels for application logs (
TRACEtoFATAL) and system logs, so production can run atWARNwithout code changes. - Infrequent Access log class: 50% lower ingestion price than Standard, for logs you keep but rarely query.
- S3 or Firehose delivery: cheaper, but it uses the Delivery log class, which loses Logs Insights and Live Tail.
30. Which loops does recursive loop detection stop, and which does it miss?
It stops loops between Lambda functions, SQS, S3, SNS and EventBridge custom buses, and it misses loops through any other service, DynamoDB included.
Per the recursive loop docs, Lambda tags events with X-Ray trace header metadata (no active tracing needed).
After about 16 invocations in one chain, it stops the invoke and sends the event to the DLQ or on-failure destination if one exists.
The scope has grown: S3 was added on October 9, 2024, and detection reached all commercial Regions on August 31, 2026. Three gaps a senior candidate should name:
- The function must call the next service with a supported SDK version, for example the JavaScript SDK 3.105.0 or later. A hand-rolled HTTP call drops the metadata.
- A DynamoDB stream that triggers a function that writes back to the same table is not caught. Budgets or Cost Anomaly Detection are the backstop.
- With an SQS source, Lambda drops the event but SQS can still redeliver the message up to the queue's
maxReceiveCount.
Intentional recursion, such as a function that pages through work by re-queueing itself, can opt out per function with PutFunctionRecursionConfig.
Signs of a Strong Answer
- They do the concurrency math out loud (requests per second times duration) before naming a setting.
- They set
MaximumRetryAttemptsandMaximumRecordAgeInSecondson every stream mapping, and can explain the one-week default. - They know where each failed event ends up for sync, async and polled sources, and what the record there contains.
- They treat
/tmpand globals as a cache that can vanish or hold leftovers, never as state. - They check whether an event filter on SQS deletes messages another consumer needs.
- They read the
REPORTandINIT_REPORTlines before guessing at a cold start or timeout cause.
Hiring Lambda Engineers
Engineers who have run Lambda at real concurrency, with streams, queues and on-call alarms, are hard to find through job boards.
Second Talent matches companies with pre-vetted DevOps and cloud engineers and back-end developers from Asia, screened with questions like these.
Tell us the stack and we send a shortlist within 24 hours. Start hiring, or use the broader AWS interview guide for the earlier rounds.






