Question: A developer submits a pull request containing 300 new UI tests for a feature. The tests are technically correct but increase CI execution time by 20 minutes. Would you approve them?
Answer:
I would review the test distribution across layers.
I would ask:
Why do we need 300 UI tests?
Which scenarios are:
Unit?
Component?
API?
Contract?
E2E?
Maybe:
300 UI tests
→ 40 critical E2E
→ 100 API
→ 100 component
→ 60 unit
would provide faster feedback and lower maintenance.
I would also evaluate:
Business risk
Critical paths
Negative cases
Browser-specific behavior
Cross-service integration
Test execution cost
Important
I would not reject tests simply because they increase execution time.
If they protect a critical payment workflow, the cost may be justified.
The decision should be:
Risk reduction
vs
Execution/maintenance cost
Question: Your framework works perfectly for 100 tests but becomes unstable at 10,000 tests. How would you determine whether the architecture itself is not scalable?
Answer:
I would look for non-linear growth.
For example:
100 tests → 2 minutes
1,000 tests → 25 minutes
10,000 → 8 hours
That suggests something is scaling poorly.
I would investigate:
Memory
Does the runner retain:
Pages
Drivers
Responses
Logs
Screenshots
Test results
File system
Are we generating millions of artifacts?
Database
Is every test creating large amounts of test data?
Reporting
Is the reporter loading all results into memory?
Logging
Is excessive synchronous logging becoming a bottleneck?
Concurrency
Does increasing workers cause:
CPU saturation
DB contention
network saturation
browser crashes
Architecture
I would measure:
Resource consumption per test
Resource consumption per worker
Resource consumption per shard
Then identify the scaling bottleneck.
Senior-level principle
A framework is scalable only when its resource consumption and execution behavior remain predictable as test volume increases.
Question: You are asked to design an automation platform from scratch for a company with 50,000 tests, 20 product teams, multiple browsers, microservices, and several CI pipelines. What architecture would you propose?
Answer: I would design it as a platform rather than one test framework.
CI/CD
|
Test Orchestrator
|
+------------+------------+
| |
Test Scheduler Test Selection
| |
+-----+------+ Risk/Tag Selection
| |
Workers Workers
| |
+----+------------+----+
| | |
UI API Integration
| | |
Playwright REST Services
Selenium Clients Kafka/etc.
|
Browsers
Supporting services
Automation Platform
|
+---------------+----------------+
| | |
Test Data Environment Artifact Store
Service Provisioning
| | |
+---------------+----------------+
|
Observability
|
+---------+---------+
| | |
Logs Metrics Traces
Test execution
I would support:
PR
↓
Fast tests
Merge
↓
Integration tests
Deployment
↓
Smoke
Nightly
↓
Extended regression
Release
↓
Risk-based regression
Production
↓
Synthetic monitoring / validation
Test selection
I would avoid running 50,000 tests for every commit.
Use:
Changed component
↓
Dependency graph
↓
Affected tests
↓
Risk-based selection
Then maintain a full regression suite separately.
Test isolation
Every execution should have controlled:
Browser context
Test data
Environment configuration
Credentials
Correlation ID
Playwright's current architecture uses isolated browser contexts per test and supports parallel workers/sharding; these are useful principles when designing large-scale execution.
Distributed execution
For Selenium-based execution, I would use a Grid-style architecture where sessions are distributed to available browser slots. Selenium Grid's current architecture separates routing, session queuing, distribution and browser nodes.
Result aggregation
Every execution should produce:
Test result
|
+-- Duration
+-- Environment
+-- Browser
+-- Worker
+-- Shard
+-- Commit
+-- Test data
+-- Failure reason
+-- Artifacts
Then provide dashboards for:
Pass rate
Initial pass rate
Flake rate
Execution time
Failure trends
Top failing tests
Top unstable environments
Browser failures
Team ownership
Failure handling
I would classify failures automatically:
PRODUCT_DEFECT
TEST_DEFECT
ENVIRONMENT
INFRASTRUCTURE
DATA
EXTERNAL_DEPENDENCY
FLAKY
UNKNOWN
That prevents engineers from spending hours investigating a browser infrastructure failure as if it were a product defect.
Governance
The platform should also enforce:
Coding standards
Test naming
Ownership
Test tagging
Retry policy
Timeout policy
Artifact policy
Dependency versions
Framework compatibility
Security rules
Most important architectural principle
I would design the platform around:
Fast Feedback
+
Reliability
+
Scalability
+
Observability
+
Test Maintainability
+
Risk Coverage
rather than around a particular automation tool.
Question: Your parallel automation framework uses a shared HashMap to store test execution results. Occasionally, results disappear or become corrupted when 100 tests run simultaneously. How would you diagnose and fix it?
Answer: Scenario
You have:
private static Map<String, TestResult> results = new HashMap<>();
Multiple test threads execute:
results.put(testId, result);
After execution, the report sometimes contains fewer results than the number of tests executed.
Interview Question
What is happening, and how would you redesign this?
What the interviewer is testing
Java collections
Thread safety
Race conditions
Concurrent collections
Parallel test architecture
Understanding versus blindly replacing HashMap
Detailed Answer
HashMap is not designed for concurrent modification by multiple threads.
The problem is not simply:
"HashMap is bad."
The real problem is:
Multiple threads are mutating shared mutable state without an appropriate synchronization strategy.
I would first establish whether the map truly needs to be shared.
My preferred solution would often be to reduce shared state, rather than immediately changing:
HashMap
to:
ConcurrentHashMap
Option 1 — Eliminate shared mutable state
If each worker can maintain its own results:
Worker 1 → results
Worker 2 → results
Worker 3 → results
↓
Aggregation
This is often easier to reason about.
Option 2 — ConcurrentHashMap
If shared access is genuinely required:
private final ConcurrentMap<String, TestResult> results =
new ConcurrentHashMap<>();
ConcurrentHashMap is designed for concurrent access.
Important
Even ConcurrentHashMap does not automatically make compound operations safe.
For example:
if (!map.containsKey(id)) {
map.put(id, result);
}
contains a race.
Another thread can insert the value between the two operations.
Instead, use atomic operations where appropriate:
map.putIfAbsent(id, result);
Senior-level answer
I would first ask:
Why is test execution sharing mutable state at all?
Then choose:
No shared state
↓
Best option
Otherwise
↓
Concurrent collection
Otherwise
↓
Explicit synchronization / atomic operation
22. A test-data allocator works perfectly with one thread but occasionally gives the same customer ID to two parallel tests. How would you fix it?
Scenario
You have:
public String getCustomerId() {
return "CUST-" + counter++;
}
Two threads execute it simultaneously.
You expect:
CUST-1001
CUST-1002
CUST-1003
but occasionally get:
CUST-1001
CUST-1001
Detailed Answer
counter++ is not an atomic operation.
Conceptually it is:
read counter
+
increment counter
+
write counter
Two threads can interleave:
Thread A → read 1001
Thread B → read 1001
Thread A → write 1002
Thread B → write 1002
Both threads may return the same original value.
Option 1 — AtomicInteger
private final AtomicInteger counter = new AtomicInteger(1000);
public String getCustomerId() {
return "CUST-" + counter.incrementAndGet();
}
Option 2 — UUID
If the database/business rules allow it:
String id = "CUST-" + UUID.randomUUID();
This can eliminate the centralized counter completely.
Option 3 — Database-generated IDs
For persistent entities, let the database generate the identifier where appropriate.
Senior-level consideration
I would ask whether uniqueness must be:
unique within JVM
unique within test run
unique within environment
globally unique
The solution depends on that requirement.
23. Your automation framework uses ThreadLocal<WebDriver>. Tests are now parallel, but memory usage keeps increasing after every suite. What would you investigate?
Detailed Answer
ThreadLocal can be useful for associating state with the current thread, but it does not automatically clean up the object.
For example:
private static ThreadLocal<WebDriver> driver =
new ThreadLocal<>();
If the thread remains alive in a thread pool, the thread-local value can remain associated with that thread unless removed.
I would investigate:
Test execution
↓
Thread pool
↓
ThreadLocal
↓
WebDriver
↓
Browser/session
Correct lifecycle
At the end of the test/worker lifecycle:
try {
// test
} finally {
WebDriver driver = threadLocalDriver.get();
if (driver != null) {
driver.quit();
}
threadLocalDriver.remove();
}
But I would also ask:
Why are we using ThreadLocal?
Modern test frameworks may already provide isolation/lifecycle management.
For example, Playwright's model uses isolated browser contexts and test fixtures rather than requiring users to build their own global thread-local browser architecture.
O
Oracle Docs
Senior-level answer
Don't use ThreadLocal simply because:
"We need parallel execution."
First understand the lifecycle and ownership of the resource.
24. Your test framework creates a fixed thread pool of 100 threads. The application only supports 20 database connections. Tests become slower as you increase threads. Why?
Scenario
You have:
ExecutorService executor =
Executors.newFixedThreadPool(100);
But:
DB connection pool = 20
and every test needs a database connection.
Detailed Answer
The system has a bottleneck.
100 test threads
↓
20 DB connections
↓
80 threads waiting
Increasing test threads does not increase database capacity.
In fact, it can make the situation worse through:
Queueing
Context switching
Connection contention
Memory usage
Lock contention
Database overload
ExecutorService controls task execution, but the optimal number of workers must account for downstream resource limits. Java's ExecutorService supports explicit lifecycle management and task submission, but it does not automatically understand your database's capacity.
O
Oracle Docs
I would measure
Worker count
DB pool utilization
DB wait time
Query latency
CPU
Lock contention
Test throughput
Then find the saturation point.
For example:
10 workers → 100 tests/min
20 workers → 190 tests/min
40 workers → 195 tests/min
80 workers → 190 tests/min
The useful concurrency level may be around 20–40, not 100.
Senior-level answer
Concurrency should be sized based on the bottleneck resource, not simply the number of CPU cores or desired test parallelism.
25. Your team uses CompletableFuture to execute three API calls in parallel. Sometimes the test hangs indefinitely. What would you investigate?
Scenario
CompletableFuture<User> user =
getUserAsync();
CompletableFuture<Order> order =
getOrderAsync();
CompletableFuture<Payment> payment =
getPaymentAsync();
CompletableFuture.allOf(user, order, payment).join();
Occasionally the test never completes.
Detailed Answer
I would investigate whether one of the underlying futures can remain incomplete indefinitely.
CompletableFuture.allOf(...) completes when all supplied futures complete; if one never completes, the aggregate future doesn't complete normally.
O
Oracle Docs
I would check:
API timeout
Connection timeout
Executor starvation
Deadlock
Blocked thread
Unbounded queue
Missing callback
External service
I would add explicit timeouts
For example:
CompletableFuture<User> user =
getUserAsync()
.orTimeout(10, TimeUnit.SECONDS);
Java's CompletableFuture provides orTimeout to exceptionally complete a future when the timeout expires.
O
Oracle Docs
I would also avoid blindly doing:
future.join();
without understanding:
Timeout
Exception handling
Cancellation
Executor behavior
Senior-level debugging
I would capture:
Future state
Thread dump
Executor queue
API latency
Connection pool
Correlation ID
The key question is:
Which future isn't completing, and why?
26. Your framework uses CompletableFuture.supplyAsync() everywhere. Performance becomes unpredictable under heavy CI load. What could be wrong?
Detailed Answer
One important question is:
Which executor is actually executing these tasks?
If an async method doesn't specify an executor, CompletableFuture uses its default asynchronous execution facility; for standard CompletableFuture, this is generally the common pool when it has sufficient parallelism.
O
Oracle Docs
If your automation framework puts many blocking operations there:
API calls
DB calls
File operations
Browser operations
you may create contention.
Example
CompletableFuture.supplyAsync(() -> callDatabase());
If callDatabase() blocks, that task occupies an executor thread while waiting.
I would consider an explicitly sized executor appropriate for that workload:
ExecutorService ioExecutor =
Executors.newFixedThreadPool(20);
and:
CompletableFuture.supplyAsync(
() -> callDatabase(),
ioExecutor
);
Senior-level answer
CompletableFuture does not mean:
"Everything is automatically scalable."
You still need to understand:
Task type
+
Executor
+
Concurrency
+
Downstream capacity
27. You have a test utility with this code. It passes in normal execution but fails under parallel execution:
class TokenManager {
private String token;
public String getToken() {
if (token == null) {
token = generateToken();
}
return token;
}
}
What is the problem?
Detailed Answer
This is a classic race condition.
Two threads can execute:
Thread A → token == null
Thread B → token == null
Thread A → generateToken()
Thread B → generateToken()
Both may generate a token.
Whether that is actually a bug depends on the desired semantics.
If exactly one initialization is required
Use appropriate synchronization or another safe initialization pattern.
For example:
public synchronized String getToken() {
if (token == null) {
token = generateToken();
}
return token;
}
But I wouldn't automatically synchronize everything.
I would ask:
Is token generation expensive?
Can different tests use different tokens?
Does token generation have side effects?
Is sharing desirable?
Is the token thread-safe?
Does the authentication server impose rate limits?
Better architecture
If each test needs isolated authentication:
Test A → Token A
Test B → Token B
Test C → Token C
may be preferable.
Senior-level principle
Thread safety and test isolation are related but different design problems.
28. A developer uses synchronized on almost every method in the automation framework "to make it thread-safe." The framework becomes extremely slow. How would you review it?
Detailed Answer
I would reject the assumption:
"More synchronization = more thread safety."
Synchronization serializes access.
Suppose:
public synchronized void executeTest() {
...
}
If 100 tests call it:
100 tests
↓
one lock
↓
effectively sequential execution
The framework may technically be safe but operationally useless.
I would identify the shared mutable state.
For each synchronized method:
What state is protected?
Why is it shared?
Can ownership be isolated?
Can it become immutable?
Can it use a concurrent collection?
Can the critical section be reduced?
Example
Instead of:
synchronized void addResult(...) {
// 100 lines
}
I would aim for:
void addResult(...) {
// prepare result outside lock
synchronized(lock) {
// only tiny critical section
}
}
where appropriate.
Senior-level answer
Thread safety should be designed around ownership and synchronization boundaries, not achieved by putting synchronized everywhere.
29. Your CI workers occasionally deadlock when tests access two shared resources: database and file system.
Scenario
Thread A:
Lock DB
↓
Lock File
Thread B:
Lock File
↓
Lock DB
The suite hangs.
What is happening?
This is a classic deadlock.
Thread A
DB lock → waiting for File
Thread B
File lock → waiting for DB
Neither can proceed.
How would you fix it?
Establish a consistent lock ordering.
For example:
Always:
DB → File
Never:
File → DB
I would also investigate
Lock ownership
Thread dumps
Lock duration
Timeout policies
Whether both locks are actually necessary
Better design
Avoid global locks where possible.
Instead:
Test A → isolated DB/data
Test B → isolated DB/data
reduces the need for synchronization entirely.
Senior-level answer
The best deadlock prevention is often:
Remove shared mutable resources rather than adding more sophisticated locking.
30. Your automation code catches Exception everywhere and logs only "Test failed". Production failures take hours to diagnose. How would you redesign exception handling?
Detailed Answer
This is an observability and error-design problem.
Bad:
try {
executeTest();
} catch (Exception e) {
throw new RuntimeException("Test failed");
}
This can destroy valuable context.
I would preserve the original exception:
throw new TestExecutionException(
"Checkout test failed for customer " + customerId,
e
);
Then the reporting system can show:
Test
↓
Business context
↓
Original exception
↓
Root cause
I would distinguish
Assertion failure
Infrastructure failure
Timeout
Application error
Test-data error
Configuration error
rather than turning everything into:
Test failed
Also important
Don't log sensitive information.
For example:
Authorization tokens
Passwords
PII
Payment information
should be masked/redacted.
Senior-level principle
An exception should answer:
What failed, where, under what context, and what caused it?
31. A test suite uses Java Streams heavily. A developer changes:
list.stream()
.filter(...)
.map(...)
.forEach(...);
to:
list.parallelStream()
.filter(...)
.map(...)
.forEach(...);
and the suite becomes flaky. Why?
Detailed Answer
parallelStream() changes the execution model.
The operations may execute concurrently, so code that depends on:
Shared mutable state
Ordering
Thread-local context
Non-thread-safe collections
External services
Browser sessions
can break.
Example:
List<String> results = new ArrayList<>();
list.parallelStream()
.map(this::executeTest)
.forEach(results::add);
ArrayList is not safe for concurrent mutation.
Another issue
The work may be I/O-heavy.
For example:
parallelStream()
↓
100 API calls
↓
external API rate limit
Now your test suite creates its own load problem.
Senior-level answer
Parallel streams are not a generic replacement for a properly designed test execution framework.
For explicit task orchestration, an ExecutorService or another appropriate concurrency abstraction provides more control over lifecycle and execution.
O
Oracle Docs
32. Your framework runs 5,000 independent API validations. A senior developer suggests using Java virtual threads to create one thread per test. Would you approve?
Detailed Answer
Potentially—but only after understanding the workload.
Virtual threads are designed to support high-throughput concurrency, especially for tasks that spend substantial time waiting on blocking I/O. They are intended to improve scale/throughput, not make individual code execute faster.
O
Oracle Docs
For example:
5,000 API calls
↓
mostly waiting for network responses
can be a good candidate.
Java provides:
try (var executor =
Executors.newVirtualThreadPerTaskExecutor()) {
...
}
The official Java documentation specifically describes this executor as useful for creating a new virtual thread for each task.
O
Oracle Docs
But I would NOT simply say:
"Virtual threads solve parallel testing."
I would investigate:
API rate limits
DB pool size
CPU
Memory
Connection pool
External dependencies
Test environment capacity
5,000 concurrent API calls may overwhelm the system under test.
Also
If the work is CPU-intensive, virtual threads don't magically provide more CPU.
Virtual threads are primarily useful for concurrency involving waiting/blocking operations, not making CPU-bound work execute faster.
O
Oracle Docs
Senior-level answer
The question isn't:
"Can Java create 5,000 threads?"
It's:
"What concurrency level can the complete system safely support?"
33. Your test framework has a static cache of API responses to improve performance. Tests pass individually but fail when executed across multiple test classes. What would you investigate?
Detailed Answer
I would immediately investigate shared mutable state and lifecycle.
For example:
static Map<String, Response> cache;
means the cache may survive beyond an individual test.
Potential problems:
Test A
↓
stores response
Test B
↓
reads stale response
or:
Test A → modifies cache
Test B → expects empty cache
Questions I would ask
What is the cache scope?
Is it immutable?
Is it thread-safe?
When is it cleared?
Can tests influence one another?
Is the cached response environment-specific?
Is it safe to share across workers?
Better options
Use:
Test-scoped cache
Worker-scoped cache
Immutable reference data
Explicit cache lifecycle
rather than uncontrolled global state.
Senior-level principle
Performance optimizations that introduce hidden state can destroy test determinism.
34. You need to implement a test-data allocator that supports 500 parallel tests. Each test needs a unique customer number. How would you design it?
Detailed Answer
I would first define the uniqueness requirement.
Then compare approaches.
Option 1 — Atomic counter
Good for a single JVM:
AtomicLong counter = new AtomicLong();
Option 2 — UUID
Good if the system accepts arbitrary identifiers.
UUID.randomUUID()
Option 3 — Database sequence
Good if IDs are database-owned.
Option 4 — Central test-data service
For distributed execution:
Worker 1 ─┐
Worker 2 ─┤
Worker 3 ─┼──> Test Data Service
Worker 4 ─┤
Worker 5 ─┘
Option 5 — Partitioned ranges
For example:
Worker 1 → 100000–100999
Worker 2 → 101000–101999
Worker 3 → 102000–102999
This reduces coordination.
Senior-level decision
For a single JVM:
Atomic counter
may be sufficient.
For distributed CI:
Central allocator
or
partitioned ID ranges
may be more appropriate.
35. You inherit a 10-year-old Java automation framework containing Singleton, Factory, Abstract Factory, Builder, Strategy, Observer, and custom thread-management classes. The framework is difficult to maintain. Would you modernize it?
Detailed Answer
Yes—but not because the design patterns are old.
The problem is accidental complexity.
I would first map:
Pattern
↓
Actual responsibility
↓
Current value
↓
Maintenance cost
For example, a Singleton may be appropriate for immutable configuration, but a Singleton WebDriver is dangerous in parallel execution.
A Factory may be useful when multiple implementations genuinely exist.
A custom thread manager may be unnecessary if standard Java concurrency APIs provide what is required.
Java's standard concurrency APIs include ExecutorService, virtual threads, and other well-tested concurrency primitives, which can reduce the need for complicated homegrown concurrency infrastructure.
Migration strategy
Measure
↓
Identify pain points
↓
Add tests around framework behavior
↓
Refactor one area
↓
Measure improvement
↓
Remove obsolete abstraction
I would not rewrite the framework simply because it is old.
Senior-level principle
Age is not technical debt. Unnecessary complexity, poor maintainability, and inability to evolve are technical debt.
15 Real-Time / Scenario-Based Questions
36. An API returns 200 OK, but the business operation actually failed. How would you design your API automation to detect this?
Scenario
You call:
POST /orders
and receive:
200 OK
{
"status": "FAILED",
"errorCode": "PAYMENT_DECLINED"
}
A junior automation test checks only:
assertEquals(200, response.statusCode());
and reports PASS.
Question
How would you design the validation so that the test detects the actual business failure?
Detailed Answer
I would separate transport-level validation from business-level validation.
HTTP validation
↓
Status code
Headers
Content-Type
Response time
Business validation
↓
Business status
Order state
Payment state
Error code
Business rules
For this response:
HTTP = successful
Business operation = failed
Therefore, checking only HTTP status is insufficient.
I would validate:
assertEquals(200, response.statusCode());
assertEquals("SUCCESS", response.jsonPath().getString("status"));
or whatever the API contract defines.
I would also validate the side effect
For an order:
POST /orders
↓
202/200
↓
GET /orders/{id}
↓
Order = CREATED
↓
Payment = AUTHORIZED
For asynchronous systems, I may need to verify eventual state rather than expect the final state immediately.
Senior-level point
A good API test validates:
Contract
+
Business semantics
+
State transition
+
Side effects
not merely the HTTP status.
HTTP status codes communicate broad classes of request outcomes, but the application can still return domain-specific information within a successful HTTP response.
M
MDN Web Docs
37. Your POST /orders API occasionally times out after 30 seconds. You don't know whether the order was created or not. Would you retry the request?
This is a very important senior-level scenario.
The answer is:
Not blindly.
Why?
Suppose:
Client
|
| POST /orders
|
↓
Order Service
|
| creates order
↓
Database
But the response is lost:
Server → 201 Created
X
network
X
Client → timeout
The client doesn't know whether the operation succeeded.
If you simply retry:
POST /orders
POST /orders
you could create two orders.
Better solution
Use an idempotency key.
POST /orders
Idempotency-Key: 8d3a-1234
The server stores the result associated with the key.
Then:
Request 1
↓
Create order
↓
Store result against key
If the client retries:
Request 2
same Idempotency-Key
↓
Return previous result
Important distinction
HTTP POST is not inherently idempotent, while methods such as PUT and DELETE are defined as idempotent in HTTP semantics. However, an application can explicitly design a POST endpoint to be safely retryable using an idempotency mechanism.
M
MDN Web Docs
What I would test
First request → success
Retry same key → same logical result
Retry same key with different payload → reject
Retry after timeout → no duplicate order
Concurrent same-key requests → one logical operation
This is a very strong Senior SDET interview answer.
38. Your microservice depends on five downstream services. Your API test fails randomly because one downstream service is slow. How would you determine whether the defect belongs to your service or the dependency?
Scenario
Test
↓
Order Service
↓
+---- Payment
+---- Inventory
+---- Customer
+---- Shipping
+---- Notification
The test occasionally takes:
2 sec
and sometimes:
40 sec
Detailed Answer
I would introduce or consume distributed tracing/correlation IDs.
For example:
Correlation ID: ABC123
Order Service
0ms → 500ms
Payment
500ms → 700ms
Inventory
700ms → 38,000ms
Now the bottleneck is clear.
I would collect:
Request/response timing
Correlation ID
Downstream status
Timeout
Retry count
Circuit-breaker state
Dependency health
Application logs
Trace spans
Test design
I would also test the dependency independently.
Order Service contract test
+
Payment contract test
+
End-to-end integration test
This avoids relying exclusively on one massive end-to-end test.
Senior-level answer
The test should provide evidence:
"Order Service failed because Inventory took 37 seconds."
rather than:
"Order test failed."
39. An API returns different JSON fields depending on the environment. QA has 30 fields, staging has 32, and production has 35. How would you automate contract validation?
Detailed Answer
I would establish an explicit API contract.
OpenAPI provides a language-independent description of HTTP APIs that can be used by humans and tools to understand request/response structure.
O
OpenAPI Initiative Publications
For example:
OpenAPI contract
↓
Request schema
↓
Response schema
↓
Automated validation
I would distinguish:
Required fields
{
"id": "...",
"status": "...",
"createdAt": "..."
}
Optional fields
{
"discount": "..."
}
The test should not fail merely because an optional field exists.
But it should fail if:
Required field removed
Wrong data type
Invalid enum
Invalid format
Unexpected breaking change
Example
Suppose:
"id": 123
becomes:
"id": "123"
If the contract says integer, that should fail.
Senior-level approach
I would run:
Schema validation
+
Semantic validation
+
Backward compatibility validation
rather than hardcoding the entire JSON response.
40. Your company has 50 microservices. A change in Customer Service unexpectedly breaks Order Service. How would you build automation to detect this before production?
Detailed Answer
I would not rely only on end-to-end tests.
I would introduce consumer-driven contract testing where appropriate.
Example:
Customer Service
↑
|
Consumer Contract
|
Order Service
The Order Service defines what it actually expects from Customer Service.
For example:
{
"customerId": "123",
"status": "ACTIVE"
}
If Customer Service changes:
{
"customerId": 123,
"state": "ACTIVE"
}
the consumer contract should detect the breaking change.
Test layers
Unit
↓
Component
↓
Contract
↓
Service integration
↓
Limited E2E
This is much more scalable than:
50 services
×
every possible combination
×
full E2E
Senior-level point
For microservices:
Contract tests answer "Can these services still communicate correctly?"
while E2E tests answer:
"Does the complete business journey work?"
You need both, but at different proportions.
41. Your Kafka consumer receives the same event twice. The first processing creates an invoice, and the second creates another invoice. How would you test and prevent this?
Scenario
Order Service
↓
Kafka
↓
Invoice Service
Event:
{
"eventId": "EVT-123",
"orderId": "ORD-100",
"type": "ORDER_CONFIRMED"
}
The same event arrives twice.
Detailed Answer
This is a classic duplicate message / at-least-once delivery problem.
The consumer should be idempotent.
For example:
eventId = EVT-123
First:
EVT-123 → process → invoice created
Second:
EVT-123 → already processed → ignore
One established approach is to persist processed message IDs and reject duplicates, often using a uniqueness constraint.
M
microservices.io
Automation should test
Single event
Duplicate event
Duplicate after consumer restart
Duplicate concurrently
Same event with retry
Out-of-order events
Malformed event
Unknown event version
Important
I would not test only:
Kafka message received
I would verify the business side effect:
1 event
↓
1 invoice
and:
2 identical events
↓
still 1 invoice
42. An API uses pagination. The first page returns 100 records, but the database contains 10,000. How would you test pagination thoroughly?
Detailed Answer
I would test more than:
page=1
Test cases
Boundary cases
0 records
1 record
99 records
100 records
101 records
10,000 records
Pagination behavior
page=1
page=2
page=3
last page
page beyond last
Page size
size=1
size=10
size=100
size=101
size=0
negative
very large
Ordering
This is critical.
If the API returns:
page 1 → IDs 1-100
page 2 → IDs 101-200
I would verify no:
duplicates
missing records
unexpected reordering
Dynamic-data scenario
Suppose records are being inserted while pagination is happening.
Offset pagination can produce:
duplicates
missing records
depending on the implementation.
I would ask whether the API uses:
offset pagination
cursor pagination
and test according to the contract.
Senior-level answer
Pagination testing should validate completeness, uniqueness, ordering and consistency, not simply HTTP 200.
43. Your API supports filtering and sorting:
GET /orders?status=PAID&sort=createdAt&direction=DESC
How would you design a high-value automation strategy without creating thousands of tests?
Detailed Answer
I would use equivalence partitioning + pairwise/combinatorial coverage + boundary testing.
Instead of testing every combination:
10 statuses
×
5 sort fields
×
2 directions
×
10 date ranges
which can explode rapidly.
I would identify:
Valid combinations
status=PAID
sort=createdAt
direction=DESC
Invalid combinations
status=INVALID
sort=UNKNOWN
direction=SIDEWAYS
Boundaries
empty
null
maximum length
maximum page size
special characters
Interaction cases
status + date
status + sort
date + sort
multiple filters
Then use representative combinations.
Senior-level principle
Good API automation optimizes for risk coverage, not raw test count.
44. A microservice returns 500 when a downstream payment service is unavailable. The requirement says it should return a graceful response instead. How would you test resilience?
Scenario
Order Service
↓
Payment Service
X
unavailable
Expected behavior:
Order Service
↓
controlled failure
rather than:
500 + stack trace + 60-second timeout
Detailed Answer
I would simulate dependency failure.
Possible techniques:
Mock service
Service virtualization
Network fault injection
Test environment dependency toggle
Controlled HTTP failures
Then test:
Payment timeout
Payment 500
Payment 503
Connection refused
Malformed response
Slow response
Partial response
Validate
HTTP response
Error code
Error message
Timeout
No sensitive information
Database state
Order state
Retry behavior
Circuit breaker behavior
Example
Payment unavailable
↓
Order state = PAYMENT_PENDING
↓
No duplicate order
↓
Retry mechanism
↓
Eventually payment succeeds
Senior-level answer
Resilience testing isn't simply:
"Verify 500."
It validates:
What happens to the business transaction when a dependency fails?
45. Your service retries failed calls three times. Under production-like load, one dependency becomes slow and your service generates thousands of additional requests. What problem could this create?
Detailed Answer
This can create a retry storm.
Suppose:
1,000 requests
↓
dependency becomes slow
↓
each request retries 3 times
↓
4,000 requests
The dependency becomes even more overloaded.
Slow dependency
↓
Retries
↓
More traffic
↓
More latency
↓
More retries
↓
System degradation
Automation should validate
Maximum retry count
Retry delay
Exponential backoff
Jitter where appropriate
Retryable status codes
Non-retryable errors
Timeout
Circuit breaker behavior
Example
Do not necessarily retry:
400
401
403
where retry won't normally correct the request.
Retrying transient failures such as certain:
429
502
503
504
may be appropriate depending on the API's contract and system design.
Important
The test should verify request count, not just final response.
For example:
Expected:
1 initial + 2 retries = 3
Actual:
1 initial + 10 retries = defect
46. An API returns 429 Too Many Requests during your automation run. The developer says, "Just add retries." Do you agree?
Detailed Answer
Not automatically.
429 indicates that the client is being rate-limited.
I would first determine:
Why are we exceeding the limit?
Possibilities:
Excessive test parallelism
Shared credentials
Missing test environment capacity
Incorrect client behavior
Actual API rate limit
Retry storm
I would test
Normal traffic → success
Threshold reached → 429
Retry-After honored
Traffic reduced → recovery
Framework design
The test framework should support configurable throttling.
500 tests
↓
20 requests/sec
rather than blindly launching:
500 requests simultaneously
OWASP explicitly identifies unrestricted resource consumption as an API security risk, including excessive concurrent requests and lack of limits on resource consumption.
O
OWASP Foundation
Senior-level point
A test that overwhelms the test environment isn't necessarily testing the application—it may simply be testing the environment's inability to handle your test framework.
47. An API accepts:
{
"userId": "123",
"role": "USER"
}
A tester changes the request to:
{
"userId": "123",
"role": "ADMIN"
}
and the server accepts it.
What kind of issue would you investigate?
Detailed Answer
I would investigate broken object property-level authorization / mass-assignment style behavior.
The API should not blindly trust client-controlled properties such as:
role
accountStatus
isAdmin
balance
permissions
if the caller is not authorized to modify them.
OWASP's 2023 API Security Top 10 specifically identifies Broken Object Property Level Authorization as a major API risk, covering improper authorization around object properties and mass-assignment/excessive-data-exposure style problems.
O
OWASP Foundation
Automation
I would create:
Regular user
↓
attempt role=ADMIN
↓
403 / controlled rejection
Then:
Admin
↓
role change
↓
allowed
Also verify:
Response
Database
Audit log
Token/claims
Senior-level point
Authorization testing must validate what the user is allowed to do, not just whether the API is authenticated.
48. Your API endpoint is:
GET /users/{userId}/orders
User A can authenticate successfully and change:
/user/123
to:
/user/456
and see User B's orders.
What would you test and how would you automate it?
Detailed Answer
This is a classic Broken Object Level Authorization (BOLA) scenario.
OWASP identifies BOLA as API1:2023 and recommends considering object-level authorization wherever an API accesses data using an object identifier supplied by the client.
O
OWASP Foundation
Test design
Create:
User A → Order A
User B → Order B
Then:
Authenticate A
GET /users/A/orders
→ 200 + A's orders
GET /users/B/orders
→ 403 / 404 according to contract
Then verify:
No B data
No metadata leakage
No count leakage
No sensitive headers
Important
Don't test only:
HTTP 403
because:
404 Not Found
may be intentionally used to avoid revealing whether another user's resource exists.
The correct expected response depends on the application's security contract.
Senior-level approach
Build reusable authorization matrices:
Role
×
Resource owner
×
Operation
×
Expected result
For example:
Actor Resource Operation Expected
User A A's order Read Allow
User A B's order Read Deny
Admin B's order Read Allow
Support B's order Read Policy-dependent
49. Your service accepts a URL from the client and fetches that URL internally. As an SDET, what security scenarios would you add?
Scenario
POST /fetch
{
"url": "https://example.com/file"
}
The server makes the outbound request.
Detailed Answer
I would investigate SSRF — Server-Side Request Forgery.
OWASP identifies SSRF as API7:2023 and specifically calls out APIs that access client-supplied URIs without proper validation.
O
OWASP Foundation
I would test whether the application:
Allows only approved destinations
Validates schemes
Prevents unexpected redirects
Restricts internal destinations
Applies network controls
Uses timeouts
Limits response size
Test categories
Valid external URL
Invalid URL
Unsupported scheme
Redirect
Large response
Slow response
Untrusted host
Internal/private destination
Malformed URL
I would perform security testing only in an authorized test environment.
Senior-level point
The test isn't:
"Does GET work?"
It is:
"Can user-controlled input cause the service to make an unauthorized outbound request?"
50. A third-party shipping API changes its response structure unexpectedly. Your service starts failing in production. How would you prevent this?
Detailed Answer
This is an external dependency contract problem.
I would introduce:
Third-party API
↓
Contract/schema validation
↓
Adapter layer
↓
Our service
Strategies
1. Contract monitoring
Regularly validate the third-party response against the expected schema.
2. Consumer contract tests
Verify the fields our application actually consumes.
3. Stubbed integration tests
Run deterministic tests without relying on the real third-party service for every CI run.
4. Compatibility handling
For example:
Old response
New response
↓
Adapter
↓
Internal model
5. Runtime observability
Detect:
Unexpected field type
Missing field
Unexpected status
Latency increase
Error-rate increase
OWASP also highlights Unsafe Consumption of APIs as a security risk: data from third-party APIs should not automatically be trusted and should be validated, sanitized and handled with appropriate transport, authentication and timeout controls.
O
OWASP Foundation
Senior-level principle
Treat external APIs as untrusted dependencies, even when they belong to a well-known provider.
51. Your microservices use asynchronous events. The Order Service publishes ORDER_CREATED, but the Notification Service sometimes receives it before the Customer Service has committed customer data. Tests fail intermittently. How would you approach this?
Detailed Answer
This is an eventual consistency / ordering / transaction-boundary problem.
Possible sequence:
Order Service
|
+--> publish ORDER_CREATED
|
+--> database commit
If the event is published before the transaction is safely committed, the consumer may observe an inconsistent state.
I would investigate
Transaction boundary
Event publication timing
Message broker semantics
Consumer retry
Event ordering
Database isolation
Consumer behavior
Test
I would intentionally create:
Customer data delay
Order creation
Event delivery
Then verify the consumer handles the temporary inconsistency correctly.
Possible architectural solutions
Depending on the system:
Transactional outbox
can ensure the event is recorded reliably as part of the database transaction and then published asynchronously.
On the consumer side, the service should be resilient to temporary unavailability:
Event received
↓
Customer unavailable
↓
Retry/backoff
↓
Customer available
↓
Process event
Senior-level point
Don't make asynchronous systems behave like synchronous systems merely to make tests deterministic.
Instead, the tests should understand and validate the eventual consistency contract.
52. Your API test suite contains 12,000 tests, but most tests simply send requests and assert status 200. The team claims API automation coverage is 90%. Do you agree?
Detailed Answer
No.
A high test count doesn't necessarily mean high API coverage.
I would measure coverage across multiple dimensions:
Endpoint coverage
+
HTTP method coverage
+
Schema coverage
+
Business-rule coverage
+
Authorization coverage
+
Negative-path coverage
+
Boundary coverage
+
State-transition coverage
+
Failure/resilience coverage
+
Security coverage
For example:
POST /orders → 200
doesn't prove:
invalid customer
duplicate order
unauthorized user
invalid payment
concurrent request
timeout
dependency failure
malformed payload
large payload
rate limit
are handled correctly.
I would create an API coverage matrix
Area Example
Happy path Valid order
Validation Missing customer
Boundary Maximum order size
Authorization Another user's order
Authentication Expired token
Concurrency Duplicate submission
Idempotency Same request twice
Dependency Payment unavailable
Resilience Timeout
Security BOLA / property authorization
Contract Schema change
Performance High request volume
Senior-level conclusion
12,000 status-code assertions are not necessarily better than 2,000 tests that validate the actual business contract.
53. An API says an order was successfully created, but the database contains no order record. How would you investigate?
Scenario
The automation test does:
POST /orders
↓
201 Created
↓
orderId = ORD-123
But:
SELECT *
FROM orders
WHERE order_id = 'ORD-123';
returns zero rows.
Interview Question
How would you determine whether this is an application defect, database issue, asynchronous processing issue, or test-data problem?
Detailed Answer
I would not immediately conclude that the API is defective.
First, I would determine the application's persistence model.
Possible architectures:
Synchronous:
API
↓
Order Service
↓
DB INSERT
↓
201
or:
Asynchronous:
API
↓
Queue/Event
↓
Order Service
↓
DB INSERT
If the second architecture is used, immediately querying the DB may produce a false failure because of eventual consistency.
Investigation sequence
Step 1 — Capture correlation ID
Request ID = ABC123
Order ID = ORD-123
Search application logs.
Step 2 — Verify API response
Was ORD-123 actually generated by the service?
Step 3 — Check event/message
ORDER_CREATED
Was published?
Step 4 — Check consumer
Did the consumer process the event?
Step 5 — Check database transaction
Was INSERT executed?
Step 6 — Check commit/rollback
The application may have executed:
INSERT INTO orders ...
but later rolled back.
Senior-level answer
I would distinguish:
API response
↓
Message/event
↓
Business processing
↓
Database transaction
↓
Database commit
and determine where the state transition stopped.
The important point is:
Don't use an immediate database assertion for an eventually consistent workflow.
54. Two tests run in parallel and both attempt to create a customer with the same email. Both tests initially see that the email doesn't exist, then both insert it. How would you prevent this?
Scenario
Both tests execute:
SELECT COUNT(*)
FROM customer
WHERE email = 'test@example.com';
Both get:
0
Then:
INSERT INTO customer(email)
VALUES ('test@example.com');
Both succeed—or one fails unpredictably.
Detailed Answer
This is a check-then-act race condition.
The problem is:
Thread A → SELECT → doesn't exist
Thread B → SELECT → doesn't exist
Thread A → INSERT
Thread B → INSERT
The application should not rely only on:
SELECT → INSERT
Database-level protection
If email must be unique:
CREATE UNIQUE INDEX ux_customer_email
ON customer(email);
Now the database becomes the final authority.
Automation should verify
Parallel request A → success
Parallel request B → controlled duplicate response
For example:
A → 201
B → 409 Conflict
depending on the API contract.
Senior-level principle
Business invariants that must always hold should generally be enforced at the database level as well as validated in application code.
Otherwise, two concurrent requests can bypass application-level checks.
55. A test passes when executed alone but fails when the entire suite runs. The database contains data left by previous tests. How would you diagnose and solve it?
Detailed Answer
This is usually a test isolation / data lifecycle problem.
I would first determine whether tests are:
Read-only
Insert-only
Update existing records
Delete records
Then investigate:
Shared test data
Static IDs
Database cleanup
Transactions
Parallel execution
Foreign-key dependencies
Test ordering
Common anti-pattern
Test A
INSERT customer ID 100
Test B
INSERT customer ID 100
Test B fails only when Test A runs first.
Better strategy
Generate unique test data:
customer-<runId>-<testId>
rather than:
customer123
Cleanup options
Option 1 — Explicit cleanup
DELETE FROM orders WHERE test_run_id = ?;
DELETE FROM customers WHERE test_run_id = ?;
Option 2 — Transaction rollback
For suitable tests:
BEGIN
↓
test
↓
ROLLBACK
Option 3 — Dedicated test database/schema
Useful for strong isolation.
Option 4 — Database reset
Useful for specific environments but potentially expensive.
Senior-level answer
I prefer:
Unique data
+
Explicit ownership
+
Reliable cleanup
+
Minimal shared state
rather than relying on test execution order.
56. Your test validates that an order total equals the sum of line items. Occasionally the values differ by 0.01. The developer says it is a rounding issue. How would you investigate?
Scenario
API:
{
"subtotal": 99.99,
"tax": 18.00,
"total": 117.98
}
But SQL calculation gives:
117.99
Detailed Answer
I would investigate numeric representation and rounding rules before changing the assertion.
Important questions:
What data type is used?
DECIMAL?
FLOAT?
DOUBLE?
For monetary values, I would expect an explicit decimal/precision policy rather than relying on binary floating-point arithmetic.
Database validation
For example:
SELECT
SUM(quantity * unit_price)
FROM order_items
WHERE order_id = ?;
Then compare against the application's defined calculation.
I would determine:
Round per line?
Round subtotal?
Round tax?
Round final total?
Banker's rounding?
Half-up?
Currency-specific rules?
These produce different results.
Test design
Don't simply write:
assertEquals(expected, actual);
without understanding the business rule.
Instead:
Line calculations
↓
Subtotal
↓
Discount
↓
Tax
↓
Shipping
↓
Final total
Validate each important stage according to the contract.
Senior-level point
For financial data, test the calculation policy—not just the final number.
57. A production-like query suddenly takes 15 seconds instead of 200 ms. The functional result is correct. As an SDET, how would you investigate?
Detailed Answer
This becomes a database performance testing problem.
I would collect:
Execution time
Query plan
Rows examined
Rows returned
Indexes
Locks
CPU
I/O
Connection pool
Database load
I would compare the execution plan before and after the regression.
For example:
EXPLAIN
SELECT ...
FROM orders
WHERE customer_id = ?;
Potential causes:
Missing index
Changed query plan
Large data growth
Stale statistics
Lock contention
Full table scan
Bad join strategy
Connection pool starvation
Important distinction
If:
DB query = 15 sec
but:
API = 15.2 sec
the bottleneck is likely downstream.
But if:
DB query = 200 ms
API = 15 sec
I would investigate application-level problems.
Senior-level answer
I would establish:
Where is the latency?
Why did it change?
Is it data-dependent?
Is it reproducible?
What is the baseline?
rather than simply reporting:
"SQL is slow."
58. A developer adds an index to make an API query faster. Read performance improves, but insert/update performance gets worse. How would you evaluate whether the index is worthwhile?
Detailed Answer
Indexes are not free.
They can improve:
SELECT
WHERE
JOIN
ORDER BY
but add overhead to writes because the index must also be maintained.
I would measure:
Before index:
SELECT = 2 sec
INSERT = 20 ms
After index:
SELECT = 100 ms
INSERT = 80 ms
Then ask:
How frequently are reads performed?
How frequently are writes performed?
Which queries benefit?
How large is the table?
Is the index actually being used?
I would inspect the query plan
An index that exists isn't necessarily an index that the optimizer will use.
Senior-level answer
Index decisions should be based on:
Query workload
+
Execution plan
+
Read/write ratio
+
Data volume
+
Latency requirements
—not simply:
"Add an index."
59. Your API creates an order and updates inventory. Occasionally the order exists but inventory wasn't reduced. How would you determine whether transaction management is correct?
Scenario
Business requirement:
Create Order
+
Reduce Inventory
=
One atomic business operation
But occasionally:
Order → CREATED
Inventory → unchanged
Detailed Answer
I would investigate whether both operations are actually part of the same transactional boundary.
Potential implementation:
BEGIN
↓
INSERT order
↓
UPDATE inventory
↓
COMMIT
If inventory update fails:
ROLLBACK
should occur if both are in the same local transaction and the business design requires atomicity.
But in microservices:
Order Service
↓
Inventory Service
may involve two different databases.
Then a single local database transaction cannot atomically cover both services.
I would ask:
Same database?
Same transaction?
Different services?
Event-driven?
Saga?
Compensation?
Automation
Test:
Normal order
Insufficient inventory
Inventory service unavailable
Inventory update timeout
Order DB failure
Duplicate request
Concurrent orders
Then validate the resulting state.
Senior-level answer
Never assume:
"@Transactional"
means the entire business workflow is atomic.
You must understand the transaction boundary.
60. Two users attempt to purchase the last available product at exactly the same time. Both API calls return success. What database/concurrency problem would you investigate?
Scenario
Initial state:
product_id = 100
stock = 1
Requests:
User A → buy
User B → buy
Result:
A → SUCCESS
B → SUCCESS
Database:
stock = -1
or:
stock = 0
with two orders.
Detailed Answer
This is a lost-update / concurrency control problem.
A naïve implementation:
SELECT stock
FROM product
WHERE product_id = 100;
then:
UPDATE product
SET stock = stock - 1
WHERE product_id = 100;
can race.
One safer pattern
Use a conditional update:
UPDATE product
SET stock = stock - 1
WHERE product_id = ?
AND stock > 0;
Then verify affected rows:
1 row → reservation succeeded
0 rows → out of stock
This makes the condition part of the atomic database operation.
Another approach
Use appropriate row locking/transaction isolation.
Automation
Run concurrent requests:
100 threads
↓
same product
↓
stock = 1
Expected:
Exactly 1 successful purchase
Remaining requests → controlled failure
Senior-level point
Concurrency defects often cannot be discovered by sequential API tests.
You need concurrent test execution plus database-state validation.
Transaction isolation levels exist specifically to control the effects of concurrent transactions, including anomalies such as dirty reads, non-repeatable reads and phantom reads.
S
SQL.org
61. A query uses LEFT JOIN, but the API is missing customers who have no orders. The developer says the query is correct. How would you debug it?
Scenario
Expected:
Customer A → 5 orders
Customer B → 0 orders
Customer C → 2 orders
The API returns:
A
C
Customer B disappears.
Detailed Answer
I would inspect whether the query effectively converts the LEFT JOIN into an INNER JOIN through a condition in the WHERE clause.
For example:
SELECT c.id, o.id
FROM customers c
LEFT JOIN orders o
ON c.id = o.customer_id
WHERE o.status = 'PAID';
The WHERE condition removes rows where o is NULL.
A condition such as:
LEFT JOIN orders o
ON c.id = o.customer_id
AND o.status = 'PAID'
may preserve customers with no matching orders.
As an SDET
I would create explicit data:
Customer A → paid order
Customer B → no orders
Customer C → unpaid order
Then verify:
A → included
B → included according to contract
C → behavior according to filter
Senior-level lesson
Database tests should contain purpose-built data that exposes JOIN semantics, rather than random data.
62. Your application deletes a customer successfully, but the database still contains orders for that customer. Is this necessarily a defect?
Detailed Answer
Not necessarily.
I would first understand the data-retention/business model.
Possible designs:
Hard delete
Customer deleted
Orders deleted
Soft delete
Customer:
deleted = true
Orders remain.
Historical retention
Orders may legally need to remain for:
Auditing
Financial records
Reporting
Compliance
Foreign-key strategy
Potential relationships:
ON DELETE CASCADE
ON DELETE SET NULL
RESTRICT
Each has different business implications.
Test question
The correct assertion isn't:
SELECT COUNT(*) FROM orders = 0;
unless that is actually the business requirement.
Instead:
What is the lifecycle contract?
Senior-level principle
Database validation must be driven by business invariants, not assumptions about how the database "should" look.
63. A search API accepts a user-provided keyword. A security scan reports possible SQL injection. How would you validate the finding as an SDET?
Scenario
API:
GET /customers?name=<input>
The application may construct SQL dynamically.
Detailed Answer
I would first determine how the query is constructed.
Unsafe pattern:
String sql =
"SELECT * FROM customer WHERE name = '" + input + "'";
This can allow user input to alter SQL semantics.
The preferred defense is a parameterized/prepared query:
PreparedStatement ps =
connection.prepareStatement(
"SELECT * FROM customer WHERE name = ?"
);
ps.setString(1, input);
OWASP recommends prepared statements/parameterized queries as a primary defense against SQL injection.
O
OWASP Cheat Sheet Series
As an SDET I would verify:
Normal input
Special characters
Unexpected SQL-like input
Long input
Unicode
Empty input
Null
But I would not treat "the request returned an error" as proof of vulnerability.
I would determine whether:
Input was treated as data
rather than:
Input changed query semantics
Additional control
The database account should follow least privilege so that even if an injection flaw exists, the blast radius is reduced. OWASP also recommends least privilege as an additional SQL-injection defense.
O
OWASP Cheat Sheet Series
64. Your database contains 50 million rows. A test validates that every API response matches database data. The suite takes 8 hours. How would you redesign the validation?
Detailed Answer
I would not query 50 million rows for every test run.
The problem is the validation strategy.
I would use multiple layers.
Layer 1 — Targeted validation
For the records created/modified by the test:
SELECT ...
FROM orders
WHERE order_id = ?;
Layer 2 — Aggregate validation
For example:
SELECT COUNT(*)
FROM orders
WHERE created_at >= ?;
Layer 3 — Sampling
For large datasets:
Random sample
Boundary records
Recently modified records
Known edge cases
Layer 4 — Reconciliation
For critical data:
API count
vs
DB count
and:
API aggregate
vs
DB aggregate
Layer 5 — Data-quality queries
For example:
SELECT COUNT(*)
FROM orders
WHERE total < 0;
or:
SELECT customer_id
FROM orders
GROUP BY customer_id
HAVING COUNT(*) > expected_limit;
Senior-level principle
Data validation should provide high confidence without unnecessarily duplicating the entire database workload.
65. A test occasionally reads stale data immediately after an update. The application uses multiple database replicas. What would you investigate?
Scenario
API UPDATE
↓
Primary DB
↓
Replication
↓
Replica
↓
API GET
The test:
PUT /customer/123
GET /customer/123
occasionally receives the old value.
Detailed Answer
This could be replication lag.
The write goes to:
Primary
while the read may go to:
Replica
before replication catches up.
I would verify
Write DB
Read DB
Replication lag
Read-routing logic
Consistency requirements
Test strategy
If the API contract promises read-after-write consistency, then:
PUT
↓
GET
should eventually or immediately reflect the update according to the contract.
If eventual consistency is expected, the test should use a bounded polling strategy:
GET
↓
not updated
↓
wait
↓
GET
↓
updated
But I would not use arbitrary sleeps such as:
Thread.sleep(10_000);
Instead:
poll until condition
with maximum timeout
Senior-level point
The test needs to know whether the system promises:
Strong consistency
or
Eventual consistency
before deciding whether the observed behavior is a defect.
66. Your test framework directly updates database tables to prepare test data. A developer says this is faster than using APIs. Would you agree?
Detailed Answer
Sometimes—but not universally.
Direct DB setup can be extremely useful for:
Large datasets
Rare states
Complex preconditions
Performance
Data cleanup
Legacy systems
For example:
INSERT INTO customer ...
INSERT INTO order ...
INSERT INTO payment ...
may be much faster than creating everything through APIs.
But there are risks.
Problem 1 — Bypassing business rules
The API might normally enforce:
Validation
Events
Audit records
Derived fields
Direct SQL bypasses them.
Problem 2 — Coupling
Tests become tightly coupled to:
Table names
Columns
Schema structure
Problem 3 — Invalid test state
You might create a database state that the real application could never create.
My preferred strategy
API/UI
↓
Normal business setup
DB
↓
Only when controlled setup is justified
For example:
Create customer via API
Create 10,000 historical orders directly in DB
Run reporting test
Senior-level answer
Use the lowest-level setup mechanism that gives reliable, fast, valid test data—but understand what business behavior you bypass.
67. A database migration changes:
customer.status
from:
VARCHAR
to:
ENUM
The application tests pass, but production deployment fails for existing records. How would you test database migrations?
Detailed Answer
This is a schema migration / backward compatibility problem.
I would test the migration against realistic pre-migration data, not an empty database.
Before migration
Create:
ACTIVE
INACTIVE
NULL
unexpected legacy values
large datasets
Then execute the migration.
Validate
Schema
Existing rows
Constraints
Indexes
Foreign keys
Application compatibility
Rollback strategy
Important scenario
Suppose existing data contains:
status = 'SUSPENDED'
but the new enum allows only:
ACTIVE
INACTIVE
The migration may fail or corrupt/lose data depending on implementation.
Zero-downtime consideration
For large production systems, I would consider an expand/contract migration approach:
Phase 1
Add new structure
↓
Phase 2
Application supports old + new
↓
Phase 3
Backfill data
↓
Phase 4
Switch reads/writes
↓
Phase 5
Remove old structure
Automation pipeline
I would run:
Fresh DB migration
+
Existing DB migration
+
Large-data migration
+
Rollback/recovery test
+
Application compatibility
Senior-level answer
A migration test should answer:
Can we safely move real existing data from schema N to schema N+1 without breaking the application or losing information?
For a 10+ year Senior SDET, I would expect the candidate to demonstrate five levels of thinking:
DATABASE
│
┌──────────────┼──────────────┐
↓ ↓ ↓
SQL Logic Concurrency Data
│ │ Integrity
↓ ↓ ↓
JOIN/Query Transactions Validation
│ │ │
└──────────────┼──────────────┘
↓
Application
↓
API / Service
↓
Test Strategy
The strongest candidates won't merely write SQL. They'll explain why the database state can diverge from API state, how concurrency creates defects, where transaction boundaries exist, how replication affects assertions, and which validations belong at the API versus database level.
_______________________________________________________________
Absolutely. Continuing with the same standard, here are 10 Playwright-focused real/scenario-based questions for a 10+ year Senior SDET / Lead SDET.
I’m deliberately avoiding basic questions like “What is a locator?” or “What is auto-waiting?”. These scenarios focus on framework architecture, reliability, parallelism, debugging, network control, authentication, CI, and production-scale Playwright usage.
I’ve cross-checked the Playwright-specific behavior against the current official Playwright documentation, including locators, isolation, authentication, network interception, retries, and tracing.
Category 5 — Playwright
10 Senior SDET / Lead SDET Scenario-Based Questions
68. Your Playwright tests pass locally but become flaky in CI. The failure is usually TimeoutError while clicking a button. How would you investigate?
Scenario
Locally:
100 tests → 100 passed
CI:
100 tests → 92 passed
8 flaky failures
Typical error:
Timeout 30000ms exceeded
waiting for locator("button").click()
Interview Question
Would you increase the timeout to 60 seconds? Why or why not?
Detailed Answer
No—not as the first solution.
A timeout is a symptom, not necessarily the root cause.
I would investigate:
Locator correctness
↓
Element visibility
↓
Element enabled state
↓
DOM stability
↓
Application API calls
↓
CI CPU/memory
↓
Network latency
↓
Browser version
↓
Test isolation
Playwright locators include auto-waiting and retry behavior, and actions such as click perform actionability checks before interacting with the element.
First question
Is this locator stable?
Bad:
page.locator("div:nth-child(4) > button").click();
Better:
page.getByRole('button', { name: 'Submit Order' }).click();
assuming that accessible role/name is stable.
Then inspect the trace
I would enable tracing in CI and inspect:
DOM snapshot
Screenshot
Network
Action timeline
Console
Errors
Playwright Trace Viewer is particularly useful for understanding what happened before and during a failed action.
I would also investigate CI resource contention
For example:
8 workers
+
CPU = 100%
+
Browser processes competing
The application may simply be slower because CI is overloaded.
Senior-level answer
I would not solve a synchronization problem by globally increasing timeouts.
I would identify:
What condition was Playwright waiting for, and why did that condition not become true?
69. Your team uses page.waitForTimeout(5000) throughout the framework. Tests are slow and still flaky. How would you refactor it?
Scenario
You find:
await page.waitForTimeout(5000);
await page.click('#submit');
await page.waitForTimeout(3000);
await expect(page.locator('.success')).toBeVisible();
Detailed Answer
I would remove arbitrary sleeps wherever possible.
A fixed sleep says:
"I hope the application is ready after 5 seconds."
A condition-based wait says:
"Continue when the required state actually exists."
For example:
await page.getByRole('button', { name: 'Submit' }).click();
await expect(
page.getByText('Order created successfully')
).toBeVisible();
Playwright's assertions automatically retry until the expected condition is met or the assertion timeout is reached.
For API-dependent UI
I may wait for a specific business event:
await page.waitForResponse(
response =>
response.url().includes('/orders') &&
response.request().method() === 'POST' &&
response.status() === 201
);
Then validate UI state.
Better architecture
Instead of:
Click
↓
sleep 5 sec
↓
assert
use:
Action
↓
Expected application state
↓
Assertion
Senior-level point
Synchronization should be based on application state, not elapsed time.
70. Your Playwright suite runs 500 tests in parallel. Tests occasionally modify each other's users, orders, and browser state. How would you redesign test isolation?
Scenario
Test A:
User = testuser@example.com
Test B:
User = testuser@example.com
Both run simultaneously.
Result:
Test A modifies user
↓
Test B sees modified state
↓
Flaky failure
Detailed Answer
I would investigate isolation at multiple levels.
Browser isolation
↓
Context isolation
↓
Authentication isolation
↓
Test-data isolation
↓
Environment isolation
Playwright creates isolated browser contexts for tests, which helps prevent cookies, local storage, and other browser state from leaking between tests.
But browser-context isolation does not isolate backend data.
That's the critical Senior SDET point.
You can have:
Browser Context A
↓
customerId = 123
and:
Browser Context B
↓
customerId = 123
Both are browser-isolated but still share the same backend customer.
Better approach
Generate test-specific data:
const user = `sdet_${testInfo.workerIndex}_${testInfo.testId}@example.com`;
or obtain data from a dedicated test-data service.
Architecture
Test 1
↓
Browser Context 1
↓
User A
↓
Order A
Test 2
↓
Browser Context 2
↓
User B
↓
Order B
Senior-level answer
Browser isolation and business-data isolation are separate concerns.
71. Your login operation takes 5 seconds and 1,000 tests need authentication. Running login through the UI for every test makes CI extremely slow. How would you optimize it without compromising isolation?
Detailed Answer
I would use Playwright's authentication-state mechanism.
Instead of:
Test
↓
Open login page
↓
Enter username
↓
Enter password
↓
Wait
↓
Application dashboard
for every test, I can establish authenticated state once and reuse it where appropriate.
Playwright documents using authenticated browser state, including storageState, to avoid repeating login steps.
Example concept:
await page.context().storageState({
path: 'playwright/.auth/user.json'
});
Then:
use: {
storageState: 'playwright/.auth/user.json'
}
But there is a major Senior-level caveat
I would not blindly share one authenticated account across 1,000 parallel tests.
If tests mutate:
profile
cart
orders
permissions
preferences
they can interfere.
Better model
Depending on application behavior:
Read-only tests
↓
Shared authenticated state may be acceptable
State-changing tests
↓
Worker/test-specific accounts
Another consideration
Authentication state can contain sensitive credentials/tokens, so it should be stored securely and excluded from source control as recommended by Playwright's authentication guidance.
Senior-level answer
The optimization isn't:
"Login once."
It is:
"Reuse authentication safely while preserving test isolation."
72. Your application uses WebSockets for real-time order updates. The UI sometimes displays the old status even though the backend has already changed it. How would you automate this reliably?
Scenario
Initial state:
Order = PROCESSING
Backend changes:
PROCESSING
↓
SHIPPED
UI receives the update through WebSocket.
Detailed Answer
I would avoid:
await page.waitForTimeout(3000);
expect(status).toHaveText('SHIPPED');
Instead, assert the eventual UI state:
await expect(
page.getByTestId('order-status')
).toHaveText('SHIPPED', {
timeout: 15000
});
But I would also validate the source event
For critical tests:
Backend
↓
Order status changed
↓
WebSocket event
↓
Browser
↓
UI state
I would investigate:
WebSocket connection
Event payload
Event ordering
Reconnect behavior
Duplicate events
Browser console errors
Important scenario
If the event is lost:
Backend → SHIPPED
X
WebSocket
does the UI recover through:
reconnection
polling
refresh
?
That becomes an important resilience test.
Senior-level answer
For real-time applications, I validate the complete state propagation path, not simply the final DOM.
73. A test needs to validate that clicking "Place Order" sends exactly one POST request, but the application automatically retries failed requests. How would you test this in Playwright?
Detailed Answer
I would use Playwright's network observation capabilities.
Conceptually:
const requestPromise = page.waitForRequest(
request =>
request.url().includes('/orders') &&
request.method() === 'POST'
);
await page.getByRole('button', { name: 'Place Order' }).click();
const request = await requestPromise;
Then inspect:
URL
Method
Headers
Payload
I would also capture the corresponding response.
But the important Senior-level scenario is retries.
Suppose:
POST #1 → 503
POST #2 → 503
POST #3 → 201
The test should verify whether:
Retry count = expected
and whether the retries are safe.
For a business operation
I would also verify:
3 HTTP attempts
≠
3 orders
The backend should maintain idempotency if retrying the operation is expected to be safe.
Senior-level answer
Don't validate only:
"POST happened."
Validate:
Request count
+
Request payload
+
Response sequence
+
Retry behavior
+
Final business state
74. Your Playwright tests use CSS selectors tied to React-generated classes. After a frontend deployment, 40% of the tests fail even though the UI behavior hasn't changed. How would you prevent this?
Scenario
Tests contain:
page.locator('.css-1a2b3c').click();
Frontend deployment changes generated CSS classes.
Tests fail.
Detailed Answer
This is a locator strategy problem, not an application defect.
I would prefer user-facing or semantic locators.
For example:
page.getByRole('button', { name: 'Submit Order' });
or:
page.getByLabel('Email');
or:
page.getByText('Order confirmed');
where appropriate.
Playwright recommends resilient locators that are tied to user-facing attributes and explicit contracts rather than brittle implementation details.
For complex applications
I would establish a locator hierarchy:
1. getByRole()
2. getByLabel()
3. getByPlaceholder()
4. getByText()
5. getByTestId()
6. CSS/XPath only when justified
data-testid
For elements without good user-facing semantics:
<button data-testid="submit-order">
then:
page.getByTestId('submit-order');
Senior-level architecture
Centralize important locators in page/component objects or reusable component abstractions where that improves maintainability—but don't hide everything behind a huge abstraction layer.
Senior-level principle
Test selectors should describe the application's stable contract, not its current DOM implementation.
75. Your Playwright test suite runs with 20 workers. Increasing workers from 10 to 20 makes the suite slower and causes more failures. How would you diagnose this?
Detailed Answer
I would not assume:
More workers = faster execution
There is a system-wide concurrency limit.
I would measure:
Browser CPU
Memory
Application CPU
Database connections
API rate limits
Network
CI machine capacity
Test-data service
For example:
Workers Runtime
--------- --------
5 40 min
10 23 min
15 18 min
20 21 min
30 30 min
The optimal point may be around 15 workers.
I would identify the bottleneck
Playwright workers
↓
Browser processes
↓
Application
↓
Database
↓
External services
Increasing concurrency at the top can overload something downstream.
Example
20 workers
↓
20 simultaneous logins
↓
Authentication API limit = 10/sec
↓
429 responses
↓
Retries
↓
More traffic
↓
Slower suite
Senior-level answer
Parallelism should be measured and tuned, not maximized blindly.
76. A Playwright test fails only once every 100 runs. The failure disappears when you run it locally with debugging enabled. How would you investigate this race condition?
Detailed Answer
This is a classic flaky-test investigation.
I would avoid immediately adding:
await page.waitForTimeout(5000);
because debugging can change timing and hide the race.
First, capture evidence
Enable:
Trace
Screenshots
Video where useful
Console logs
Network
Test metadata
Browser logs
Playwright's trace functionality is specifically designed to help inspect test execution after failures, including actions, snapshots and network activity.
Then classify the race
Potential patterns:
UI rendered before data
Data arrived before listener attached
Two API calls completed out of order
WebSocket event arrived early
Test cleanup raced with next test
Shared backend data changed
Multiple browser tabs/windows
Example race
Bad:
await page.click('#submit');
page.on('response', handler);
The response may occur before the listener is registered.
Better pattern:
const responsePromise = page.waitForResponse(...);
await page.click('#submit');
const response = await responsePromise;
Retry usage
If Playwright retries a failed test, I would use the retry result for diagnosis—not as proof that the test is healthy.
A test that passes on retry is still a flaky test.
Senior-level answer
A retry that turns red into green doesn't fix the race condition; it only makes the symptom less visible.
77. Your Playwright suite has UI tests that mock almost every API response. CI is green, but production frequently breaks because of backend changes. How would you redesign the test strategy?
Detailed Answer
This is a test realism problem.
Mocking is valuable, but excessive mocking can produce:
Test application
↓
Fake API
↓
PASS
Production
↓
Real API
↓
FAIL
I would use a layered strategy
UI
│
┌──────────┼──────────┐
↓ ↓ ↓
Mocked Real API E2E
tests tests
│ │ │
Fast Medium Slow
Stable Realistic Broad
Mock when testing:
Error states
Rare backend responses
Network failures
Slow responses
Specific edge cases
Unstable third-party services
Use real services when testing:
Critical user journeys
API contracts
Authentication
Business workflows
Integration behavior
Add contract testing
For example:
Frontend expectation
↓
API contract
↓
Backend implementation
This catches API schema changes earlier.
Senior-level principle
Mocking should isolate a behavior under test—not eliminate the dependencies that define whether the system actually works.
78. Your regression suite passes locally but fails in CI because different tests fail on different runs. How would you determine whether the problem is the test, environment, or pipeline?
Scenario
Local:
500 tests → 500 passed
CI:
Run #1 → Test A, Test F failed
Run #2 → Test B, Test F failed
Run #3 → Test C, Test H failed
There is no consistent failure.
Interview Question
As a Lead SDET, how would you systematically isolate the problem?
Detailed Answer
I would treat this as an environment/pipeline reliability investigation, not immediately as 20 independent test defects.
I would divide the investigation into layers:
Test
↓
Browser/runtime
↓
Container/VM
↓
Application
↓
Database
↓
External dependencies
↓
CI infrastructure
Step 1 — Identify failure patterns
I would collect:
Test name
Worker
Node/container
Browser version
OS
Commit
Environment
Failure type
Execution duration
Retry result
For example:
Test A → Worker 4 → container-17 → timeout
Test B → Worker 2 → container-03 → DB connection
Test C → Worker 4 → container-17 → timeout
If multiple unrelated tests fail on the same worker/container, I would investigate that infrastructure.
Step 2 — Compare local vs CI
Check:
Node/Java version
Browser version
Environment variables
Timezone
Locale
CPU
Memory
Network
Database
Secrets
Dependency versions
Step 3 — Check resource saturation
For example:
CPU = 100%
Memory = 95%
A browser timeout may actually be caused by resource starvation.
Step 4 — Re-run the exact CI artifact
I would reproduce using the same:
Docker image
Browser
Test commit
Environment variables
Configuration
Step 5 — Look for test-order dependency
Run:
test A → test B
versus:
test B → test A
and then run the suite in randomized order.
Senior-level answer
I would build a failure fingerprint rather than simply rerunning the failed tests.
The key question is:
Does the failure follow the test, the environment, the worker, the test order, or the infrastructure?
79. Your pipeline currently runs 1,500 automated tests sequentially and takes 3 hours. Management wants it reduced to 30 minutes. How would you redesign the pipeline?
Scenario
Current:
Build
↓
1,500 tests
↓
3 hours
↓
Deploy
Target:
< 30 minutes
Detailed Answer
I would not simply increase CI agents from 1 to 20.
First, I would measure test duration.
Example:
Test suite = 180 min
UI = 120 min
API = 40 min
DB = 15 min
Other = 5 min
Then parallelize based on test characteristics.
Possible architecture
Build
│
┌───────────┼───────────┐
↓ ↓ ↓
API Tests UI Tests DB Tests
│ │ │
5 workers 15 workers 3 workers
But parallelization requires isolation
I would verify:
Unique test data
Independent users
Independent browser contexts
Database isolation
External dependency limits
No test ordering
Optimize test distribution
If worker 1 receives:
10 tests × 5 min = 50 min
while worker 2 receives:
30 tests × 30 sec = 15 min
the pipeline is poorly balanced.
I would use historical execution duration to distribute tests.
Additional optimization
Move tests to the appropriate layer:
UI E2E
↓
Critical workflows only
API
↓
Most business logic
Unit/component
↓
Large-volume validation
Senior-level answer
The goal isn't:
"Run more tests simultaneously."
The goal is:
Reduce feedback time while preserving confidence, determinism, and coverage.
80. A developer pushes a commit. Unit tests pass, but your integration tests fail because the database schema is incompatible. How would you prevent this from reaching the main branch?
Detailed Answer
I would implement a quality gate before merge.
Pipeline:
Commit
↓
Compile
↓
Unit Tests
↓
Static Analysis
↓
Build Artifact
↓
Database Migration
↓
Integration Tests
↓
API Tests
↓
Critical E2E
↓
Merge
The important part is that integration tests run against the same migration path that production uses.
Example
The migration changes:
customer.status
but the application still expects the old value.
The CI environment should create or upgrade a database using the actual migration scripts.
I would test both:
Empty DB
and:
Existing DB + realistic data
because a migration can succeed on an empty database while failing against production data.
Merge protection
The main branch should require:
Required checks = PASS
before merging.
Senior-level answer
Schema changes must be tested as part of the application delivery pipeline—not as a separate DBA activity.
81. Your production deployment succeeds, but the application immediately starts returning 500 errors. The deployment pipeline says "SUCCESS." What is missing?
Scenario
Pipeline:
Build → PASS
Tests → PASS
Deploy → PASS
Production:
HTTP 500
HTTP 500
HTTP 500
Detailed Answer
A successful deployment only proves that the deployment mechanism completed.
It does not prove that the application is healthy.
I would add post-deployment validation.
Pipeline
Deploy
↓
Health Check
↓
Smoke Tests
↓
Critical API Tests
↓
Monitoring Validation
↓
Release
Health check
For example:
GET /health
But a simple health endpoint may not be sufficient.
I would validate:
Application startup
Database connectivity
Required dependencies
Authentication
Critical API
Critical UI workflow
Canary deployment
For high-risk changes:
Deploy 5%
↓
Validate
↓
Deploy 25%
↓
Validate
↓
Deploy 100%
Automatic rollback
If:
Error rate > threshold
then:
Rollback
Senior-level answer
A deployment pipeline should answer:
"Is the new version actually working in the target environment?"
—not merely:
"Did Kubernetes/Jenkins/etc. report deployment success?"
82. Your CI pipeline randomly fails because an external payment API returns 503. Should the pipeline retry the test, mock the API, or fail the build?
Detailed Answer
There isn't one universal answer.
I would classify the test.
For a true end-to-end payment test
A real 503 may be meaningful.
I would fail the test if the purpose is:
Validate payment-provider integration
and record:
External dependency unavailable
For application business-logic tests
I would mock the provider:
Application
↓
Payment interface
↓
Mock
Then explicitly test:
200
400
401
402
429
500
503
timeout
For transient infrastructure failures
A limited retry may be appropriate.
But:
Retry 5 times
is not a substitute for fixing instability.
Better strategy
Separate:
Application tests
+
Contract tests
+
Provider integration tests
+
End-to-end tests
Senior-level principle
Retries should handle genuinely transient failures; mocking should isolate dependencies; real integration tests should verify real integrations.
Don't use one mechanism to solve all three problems.
83. Your CI pipeline stores username/password/API tokens directly in the YAML file. A security review flags it. How would you redesign the pipeline?
Scenario
Current:
env:
DB_PASSWORD: "MyPassword123"
API_TOKEN: "abc123..."
Detailed Answer
Credentials should not be committed to source control.
I would move secrets into a proper secret-management mechanism, such as the CI platform's secret store or an enterprise secret manager.
Pipeline becomes conceptually:
CI Pipeline
↓
Secret Manager
↓
Runtime injection
↓
Application/Test
Important controls
Secrets should be:
Encrypted at rest
Masked in logs
Scoped appropriately
Rotated
Short-lived where possible
Least privilege
Unavailable to untrusted jobs
Also inspect logs
Even if the secret isn't in YAML, this is dangerous:
echo $API_TOKEN
or:
curl ...?token=secret
because CI logs may expose credentials.
Pull-request security
I would be especially careful with untrusted PRs because running arbitrary code with production-capable secrets can create a major security risk.
Senior-level answer
Secrets should be injected at runtime, tightly scoped, masked, rotated, and never treated as normal configuration.
84. Your organization has 10 microservices. A change in Service A causes 200 unrelated E2E tests to run. The pipeline takes 90 minutes. How would you improve the CI/CD strategy?
Detailed Answer
I would introduce change-aware test selection, but carefully.
For example:
Change:
payment-service
Potentially run:
Payment unit tests
Payment integration tests
Payment contract tests
Affected API tests
Critical cross-service tests
instead of every UI test.
Dependency graph
I would build:
Frontend
↓
Order Service
↓
Payment Service
↓
Payment DB
and:
Frontend
↓
Catalog Service
↓
Catalog DB
If only Catalog changes, payment-specific tests don't necessarily need to run for every commit.
But I would maintain multiple gates
Pull request
Fast tests
+
Affected tests
+
Contract tests
Main branch
Full regression
Nightly
Large-scale
cross-service
full E2E
performance
resilience
Senior-level caution
Test selection must not become a blind optimization.
You need confidence that the dependency graph is accurate.
Senior-level answer
Optimize feedback using test impact analysis while retaining scheduled full regression for coverage protection.
85. Your Docker-based test containers work on one CI agent but fail on another. The error is "works on my machine" all over again. How would you solve it?
Detailed Answer
I would eliminate environmental drift.
First, capture:
Docker version
OS/kernel
CPU architecture
Container image
Browser version
Node/Java version
Environment variables
Mounted volumes
Network configuration
Use immutable images
Instead of:
node:latest
prefer a controlled version:
node:<specific-version>
or a company-maintained image.
Similarly, browser versions should be controlled.
Build once, run consistently
Pipeline:
Build image
↓
Tag immutable image
↓
Run tests
↓
Publish artifact
rather than rebuilding different environments for different stages.
Container image
Conceptually:
Base image
+
Runtime
+
Browser
+
Dependencies
+
Test framework
+
Known configuration
Senior-level answer
The objective is:
Same source + same artifact + same runtime = reproducible execution.
86. Your pipeline passes all tests, but the deployment artifact is accidentally built from a different commit than the one tested. How would you prevent this?
Scenario
Pipeline:
Commit A
↓
Tests PASS
Later:
Commit B
↓
Build artifact
↓
Deploy
Now production contains code that was never tested.
Detailed Answer
This is a serious artifact traceability problem.
I would use immutable artifacts.
Pipeline:
Source Commit
↓
Build
↓
Artifact
↓
Test EXACT artifact
↓
Promote SAME artifact
↓
Production
Not:
Build
↓
Test source
↓
Rebuild
↓
Deploy
because the rebuild could produce a different artifact.
Artifact metadata
Every artifact should be traceable to:
Git SHA
Build ID
Version
Dependencies
Build timestamp
For example:
order-service:1.8.4
commit=abc123
build=7845
Promotion model
Artifact
↓
DEV
↓
QA
↓
STAGING
↓
PRODUCTION
The same artifact is promoted.
Senior-level answer
Never rebuild between validation and production promotion when you can promote the exact tested artifact.
87. Your CI pipeline retries failed tests automatically three times. The dashboard reports 99.8% pass rate, but the team later discovers many flaky tests. How would you prevent retries from hiding quality problems?
Detailed Answer
This is a very common mature-CI problem.
Suppose:
1,000 tests
First attempt:
980 PASS
20 FAIL
Retry:
19 PASS
1 FAIL
Dashboard reports:
999 PASS
1 FAIL
But the real first-attempt stability is:
98%
I would track both metrics
First-pass rate
980 / 1000 = 98%
Final pass rate
999 / 1000 = 99.9%
Both are useful—but they mean different things.
Track flaky tests separately
Example:
Test First Pass Retry Pass
------------------------------------------------
CheckoutTest FAIL PASS
SearchTest PASS PASS
PaymentTest FAIL PASS
Then classify:
PASS
FAIL
FLAKY
INFRASTRUCTURE FAILURE
Retry policy
Retries should be:
Limited
Visible
Tracked
Non-zero-cost
For example:
1 retry
rather than:
10 retries until green
Quality dashboard
I would track:
First-pass rate
Final pass rate
Flaky-test rate
Mean test duration
Failure recurrence
Infrastructure failure rate
Senior-level principle
Retries are a diagnostic safety net, not a quality strategy.
If a test passes only after retry, I still consider it an engineering problem that needs investigation.
A strong Lead SDET should be able to design something closer to:
Git Commit
│
▼
┌─────────────┐
│ Build │
└──────┬──────┘
│
┌─────────┴─────────┐
▼ ▼
Fast Quality Gates Security Scan
│ │
└─────────┬─────────┘
▼
Test Exact Artifact
│
┌──────────┼──────────┐
▼ ▼ ▼
Unit API E2E
│ │ │
└──────────┼──────────┘
▼
Integration
│
▼
Deployment
│
▼
Smoke / Health Check
│
▼
Canary / Rollout
│
▼
Monitoring
│
┌──────┴──────┐
▼ ▼
Healthy Unhealthy
│ │
▼ ▼
Promote Rollback
_______________________________________________________________________________------___
Absolutely. For Leadership / System Design, I’ll raise the level further. These are aimed at someone interviewing for Senior SDET / Lead SDET / SDET Architect with 10+ years of experience.
I’m avoiding generic leadership questions like “What are your strengths?” and focusing on real situations where you have to make architectural, organizational, quality, and delivery decisions.
Category 7 — Leadership / System Design
5 Real / Scenario-Based Questions
Questions 88–92
88. You join a company where 2,000 automated tests exist, but the team doesn't trust the results. How would you take ownership and turn automation into a reliable quality system?
Scenario
You join as Lead SDET.
The team tells you:
"Automation is already there."
But you discover:
2,000 tests
↓
25% flaky
↓
Long execution time
↓
Frequent environment failures
↓
Duplicate coverage
↓
Tests regularly ignored
Developers don't wait for the automation result before merging.
Interview Question
What would your first 90 days look like?
Detailed Answer
I would not start by rewriting the entire framework.
My first objective would be to understand why the organization doesn't trust automation.
I would evaluate five dimensions:
Reliability
Speed
Coverage
Maintainability
Feedback value
Phase 1 — Assessment
I would collect:
Test count
Execution time
First-pass rate
Flaky rate
Failure categories
Duplicate tests
Code coverage
Business-risk coverage
Environment failures
Then classify failures:
TEST DEFECT
APPLICATION DEFECT
ENVIRONMENT
DATA
INFRASTRUCTURE
FLAKINESS
This is important because:
100 failures
doesn't necessarily mean:
100 product defects
Phase 2 — Stabilize
I would prioritize the highest-value failures.
For example:
Top 20 flaky tests
Top 10 infrastructure problems
Top slowest suites
Critical business-flow failures
Rather than trying to fix all 2,000 tests simultaneously.
Phase 3 — Rationalize
I would identify:
Duplicate tests
Low-value UI tests
Tests better suited to API level
Tests better suited to unit/component level
Move coverage toward:
UI
/ \
API E2E
|
Component
|
Unit
Phase 4 — Establish quality gates
For example:
PR:
Fast tests + affected tests
Main:
Broader regression
Nightly:
Full regression
Release:
Critical E2E + integration + smoke
Phase 5 — Create ownership
Every persistent flaky test should have:
Owner
Priority
Reason
Tracking ticket
Expected resolution
Metrics
I would publish:
First-pass rate
Flaky-test rate
Mean execution time
Defect detection rate
Escaped defects
Automation coverage of critical workflows
Senior/Lead-level answer
My goal isn't:
"Increase the number of automated tests."
It is:
"Create a quality signal that developers and release managers can trust."
89. Your team wants to automate 100% of regression tests through UI because "UI automation represents the real user." You disagree. How would you convince the team?
Scenario
Product has:
500 business scenarios
Management says:
"Automate all of them using Playwright."
You estimate:
500 UI tests
→ 4 hours
→ high maintenance
→ frequent UI failures
Interview Question
How would you design the automation strategy instead?
Detailed Answer
I would introduce a risk-based automation pyramid / test distribution strategy.
Not every business rule needs to be validated through the browser.
For example:
Business Rules
│
├── Unit/component
│
├── API/service
│
├── Integration
│
└── UI/E2E
Example
Suppose checkout contains:
Tax calculation
Discount calculation
Inventory rules
Payment validation
Order creation
UI confirmation
I wouldn't test all combinations through UI.
Instead:
Unit/component
Tax calculation
Discount rules
API
Order creation
Payment behavior
Inventory validation
Integration
Order → Payment → Inventory
UI
Customer adds item
↓
Checkout
↓
Place order
↓
Confirmation
Why?
Because UI tests are generally:
Slower
More expensive
More fragile
Harder to diagnose
But I would not completely eliminate UI coverage.
Critical user journeys should remain.
For example:
Login
Checkout
Payment
Order history
Critical admin workflow
How I would communicate this to leadership
Instead of saying:
"UI automation is bad."
I'd show numbers.
Example:
500 UI tests
= 4 hours
= 12% flaky
Proposed:
100 UI
250 API
100 integration
50 component
= 35 minutes
= significantly better diagnostics
Senior-level principle
Automation strategy should optimize confidence per unit of execution and maintenance cost—not maximize UI test count.
90. You are asked to design an automation architecture for a new microservices product with 20 services. How would you design it as Lead SDET?
Scenario
The platform has:
20 microservices
3 frontend applications
5 databases
Kafka/event streaming
External payment provider
External notification provider
The company wants:
Fast PR feedback
Reliable regression
Production confidence
Parallel execution
Easy debugging
Interview Question
Design the high-level test automation architecture.
Detailed Answer
I would design it around test layers and service boundaries, rather than building one giant E2E framework.
High-level architecture
CI/CD
│
┌───────────┼───────────┐
▼ ▼ ▼
Unit Service Contract
Tests Tests Tests
│ │ │
└───────────┼───────────┘
▼
Integration
│
▼
API Tests
│
▼
Critical E2E
│
▼
Production
Service-level tests
Each microservice should own:
Unit tests
Component tests
Service/API tests
Contract tests
Contract testing
This is particularly important with 20 services.
Example:
Order Service
↓
Payment Service
The Order Service depends on:
{
"paymentId": "123",
"status": "SUCCESS"
}
A contract test can detect breaking changes without requiring the entire platform E2E suite.
Integration tests
Validate real boundaries:
Service
↓
Database
or:
Order
↓
Kafka
↓
Inventory
E2E
Keep E2E focused on critical business journeys:
Login
↓
Browse
↓
Cart
↓
Checkout
↓
Payment
↓
Order confirmation
Test-data architecture
I would design:
Test Data Factory
│
├── API setup
├── DB setup where justified
└── Cleanup
with unique data per test/worker.
Execution architecture
Test Orchestrator
│
┌────────────────┼────────────────┐
▼ ▼ ▼
API workers Integration E2E workers
│ │ │
▼ ▼ ▼
Container A Container B Browser C
Observability
Every test should have:
Correlation ID
Test ID
Build ID
Environment
Service logs
Request/response metadata
Trace
Screenshot/video where useful
This makes:
Test failure
↓
Service
↓
Request
↓
Log
↓
Root cause
much easier.
Senior-level answer
The biggest mistake would be creating:
20 services
↓
1 giant E2E suite
↓
3 hours
↓
flaky
Instead:
Test each boundary at the cheapest reliable layer and reserve E2E for business-critical system behavior.
91. Two senior engineers strongly disagree about whether a critical test should be mocked or integrated with the real external service. How would you make the decision?
Scenario
Your payment provider is unreliable in the test environment.
Engineer A:
"Always use the real payment API. Otherwise the test isn't realistic."
Engineer B:
"Always mock it. Otherwise the suite is flaky."
Both are experienced engineers.
Interview Question
As Lead SDET, how would you resolve the disagreement?
Detailed Answer
I would avoid choosing based on opinion.
I'd first ask:
What behavior are we trying to prove?
There may actually be multiple tests with different purposes.
Test 1 — Application behavior
Use a mock:
Application
↓
Payment mock
Test:
Payment success
Payment declined
Timeout
503
Invalid response
This gives deterministic coverage.
Test 2 — Contract/integration
Use the real provider where practical:
Application
↓
Payment API
Validate:
Authentication
Request schema
Response schema
Protocol
Test 3 — Critical E2E
Use a controlled real integration or provider sandbox for a limited number of workflows.
Test 4 — Failure handling
Mock provider failures deliberately.
Decision matrix
Test purpose Mock Real service
Business logic ✅
Error scenarios ✅
Contract ✅
Integration ✅
Critical E2E ✅
Third-party outage simulation ✅
Performance of our service Often ✅
Leadership aspect
I would make the decision based on:
Risk
Reliability
Cost
Coverage
Execution time
Environment stability
Business criticality
rather than:
Engineer A vs Engineer B
Senior-level answer
The correct architecture is often both.
Mock to control dependencies; integrate to validate boundaries.
92. Your release team asks: "Can you guarantee that this release has zero production defects?" How would you respond as Lead SDET?
Scenario
A critical release is planned tomorrow.
Management asks:
"Can QA guarantee there will be no production bugs?"
Detailed Answer
I would not claim something that testing cannot prove.
I would explain:
No responsible engineering team can guarantee zero defects solely through testing. We can provide evidence-based confidence and quantify known risk.
Then I would provide a release-quality assessment.
Example
Critical scenarios 100% passed
High-risk scenarios 98% passed
Regression 99.2% first-pass
Known defects 2 medium
Critical defects 0
Security scan PASS
Performance PASS
Production smoke PASS
Then identify residual risk:
Known:
- Medium defect in reporting
Unknown:
- Third-party provider behavior under peak load
I would create a release risk matrix
Area Risk Evidence
Authentication Low Full regression passed
Checkout Low E2E + API + integration
Payment Medium Provider sandbox limitation
Reporting Medium Known defect
Performance Low Load test passed
Go/no-go decision
The SDET/QA role is not necessarily to say:
YES
or:
NO
without context.
Instead:
Quality Evidence
+
Known Risks
+
Business Impact
+
Release Criteria
↓
Go / No-Go decision
Leadership responsibility
If a release has:
Known critical defect
I would clearly communicate:
Impact
Probability
Affected users
Workaround
Evidence
Recommendation
Senior-level answer
A strong Lead SDET doesn't promise:
"There are no bugs."
They provide:
"Here is the evidence, here are the known risks, here is what we tested, here is what we couldn't test, and here is our confidence level."
Leadership / System Design — Complete
# Scenario What it evaluates
88 2,000 tests but nobody trusts automation Automation transformation
89 Team wants 100% UI automation Test strategy
90 Design automation for 20 microservices System/test architecture
91 Mock vs real external service disagreement Technical leadership
92 Management asks for zero-defect guarantee Risk-based leadership
What these 5 questions cover
Leadership / System Design
│
├── Automation Strategy
├── Test Pyramid
├── Risk-Based Testing
├── Framework Architecture
├── Microservices Testing
├── Contract Testing
├── Test Data Architecture
├── Distributed Test Execution
├── Observability
├── Technical Decision Making
├── Stakeholder Management
├── Quality Metrics
├── Release Risk Management
└── Engineering Leadership
______________________________________________________________________
Absolutely. Continuing with the same standard, these will be Lead SDET-level Performance / Load / Scalability scenarios, not tool-definition questions.
Category 8 — Performance / Load / Scalability Testing
10 Real / Scenario-Based Questions
Questions 93–102
These questions focus on workload modeling, bottleneck analysis, distributed systems, database performance, scalability, capacity planning, SLAs/SLOs, and production-like performance engineering.
93. Your application performs well with 1,000 users but becomes extremely slow at 10,000 users. How would you identify the bottleneck?
Scenario
Performance results:
Concurrent Users Avg Response p95 Error Rate
1,000 180 ms 300 ms 0%
3,000 250 ms 450 ms 0%
5,000 700 ms 1.8 sec 1%
10,000 4.5 sec 12 sec 8%
Interview Question
How would you determine whether the bottleneck is application code, database, infrastructure, network, or an external dependency?
Detailed Answer
I would not conclude that the application is simply "unable to handle 10,000 users."
First, I would correlate load-test results with system telemetry.
Load Test
↓
Response Time
↓
Application Metrics
↓
Database Metrics
↓
Infrastructure Metrics
↓
External Dependency Metrics
Application metrics
I would examine:
CPU
Memory
GC
Thread pools
Connection pools
Request queue
Request throughput
Error rate
Latency
For example:
CPU = 98%
Thread pool = exhausted
DB CPU = 40%
This suggests the application tier may be the bottleneck.
But:
App CPU = 40%
DB CPU = 98%
DB connections = exhausted
points toward the database.
Database investigation
I would examine:
Slow queries
Query execution plans
Indexes
Lock contention
Connection pool
Deadlocks
I/O
CPU
Cache hit ratio
External services
Suppose:
Application latency = 4 sec
Internal processing = 200 ms
Payment API = 3.5 sec
Then increasing application servers won't solve the problem.
Network
I would investigate:
Latency
Packet loss
Bandwidth
Connection establishment
TLS overhead
Important Lead SDET concept
I would correlate:
Users
↓
Requests/sec
↓
Response time
↓
CPU
↓
DB latency
↓
External API latency
rather than looking at the load-test report alone.
Senior answer
Performance testing without system observability tells you that something is slow. Performance engineering tells you why.
94. Product management says: "We expect 100,000 users next month." They ask you to design a performance test. What information do you need before creating the test?
Detailed Answer
I would not immediately create 100,000 virtual users.
"100,000 users" does not define a workload.
I need to understand actual usage patterns.
Questions I would ask
1. Are these concurrent users?
100,000 registered users
≠
100,000 concurrent users
Maybe:
100,000 daily active
15,000 peak concurrent
2. What are users doing?
For example:
Login → 10%
Search → 40%
Product → 25%
Checkout → 15%
Reports → 10%
3. What is the expected traffic pattern?
Constant
Ramp-up
Peak
Spike
Soak
4. What are the SLAs?
For example:
p95 < 500 ms
p99 < 1 sec
Error rate < 0.1%
5. What infrastructure exists?
Application instances
Database
Cache
Queue
Load balancer
External APIs
6. What is the expected growth?
Maybe:
Today → 10K
3 months → 50K
6 months → 100K
Then I build a workload model
15K peak users
│
┌───────────┼───────────┐
↓ ↓ ↓
Search Browse Checkout
40% 45% 15%
Senior-level answer
A performance test should model business workload, not merely generate a large number of virtual users.
95. Your API has an SLA of p95 < 500 ms. Average response time is only 200 ms, so the team says performance is good. You disagree. Why?
Scenario
Metrics:
Average = 200 ms
p50 = 150 ms
p95 = 1.2 sec
p99 = 5 sec
Detailed Answer
The average is hiding the tail latency.
If:
p95 = 1.2 sec
and SLA is:
p95 < 500 ms
then the system fails its SLA.
Why average is dangerous
Imagine 100 requests:
95 requests → 100 ms
5 requests → 5 seconds
The average may still look acceptable.
But those 5% of users experience terrible performance.
I would examine:
p50
p90
p95
p99
max
and correlate slow requests with:
Endpoint
User journey
Database query
Instance
Region
Payload size
External dependency
Lead-level consideration
For customer-facing systems, tail latency is often more useful than average latency.
For example:
Requirement:
p95 < 500ms
p99 < 1s
Error rate < 0.1%
The performance gate should reflect those requirements.
Senior answer
Average latency describes the center of the distribution; percentile latency tells you what a meaningful portion of users actually experience.
96. Your system passes a 30-minute load test but fails after 8 hours of continuous traffic. How would you investigate?
Scenario
30-minute test → PASS
8-hour test → FAIL
After several hours:
Memory usage → steadily increasing
GC → increasing
Response time → increasing
Errors → increasing
Detailed Answer
This strongly suggests a possible resource leak or gradual degradation.
I would run a soak/endurance test and monitor:
Memory
Heap
GC
Threads
Connections
File descriptors
DB connections
Cache
Queues
CPU
Disk
Example
Hour 1:
Memory = 2 GB
Hour 4:
Memory = 3 GB
Hour 8:
Memory = 6 GB
That pattern requires investigation.
Potential causes
Memory leak
Unreleased connections
Thread leak
Unbounded cache
Message backlog
Database connection leak
File descriptor leak
Log accumulation
Garbage collection pressure
I would compare:
Start-of-test state
vs
End-of-test state
For example:
Threads:
500 → 5,000
DB connections:
100 → 1,000
Important point
A soak test is not simply:
"Run the load longer."
It is designed to detect long-term stability problems.
Senior answer
Short load tests validate immediate capacity; endurance tests validate system stability over time.
97. Your application suddenly receives 5× normal traffic because of a marketing campaign. The system normally handles the traffic but crashes when the traffic arrives suddenly. What type of performance test would you design?
Detailed Answer
This is a spike-load scenario.
Normal:
1,000 req/sec
Spike:
1,000
↓
5,000 req/sec
within a very short period.
I would test:
Baseline
↓
Sudden spike
↓
Peak
↓
Return to baseline
What I would observe
Auto-scaling
Queue depth
Connection pools
CPU
Memory
Load balancer
Database
Cache
Error rate
Recovery time
Important question
Does the system recover?
Suppose:
Traffic spike
↓
System overloaded
↓
Traffic returns to normal
↓
System remains unhealthy
That may indicate:
Queue backlog
Connection exhaustion
Memory pressure
Failed instances
Database saturation
I would also test autoscaling
For example:
1 instance
↓
Traffic spike
↓
Auto-scale to 10
↓
Stabilize
↓
Scale down
The test should verify whether scaling happens quickly enough.
Senior answer
Spike testing validates how the system behaves under sudden traffic changes, not just sustained load.
98. Your API response time is acceptable, but the database CPU reaches 100% during load testing. The application team says, "The API is still fast, so this isn't a problem." How would you respond?
Detailed Answer
I would challenge the conclusion.
A performance problem can exist before the user-visible SLA is violated.
If:
API = 300 ms
DB CPU = 100%
the system may currently have little headroom.
I would investigate
Top SQL queries
Query frequency
Query execution plan
Indexes
Locks
Connection pool
Read/write ratio
Caching
Database scaling
Example
Suppose:
Query A:
20 ms × 20,000 executions/sec
Even though each individual query is fast, the cumulative workload can saturate the database.
Capacity question
I would ask:
What happens at 2× today's traffic?
If:
Current:
DB CPU = 100%
then there is essentially no capacity margin.
Possible solutions
Index optimization
Query optimization
Caching
Read replicas
Connection-pool tuning
Partitioning
Data-model changes
Horizontal/vertical scaling
Senior answer
Performance testing should evaluate capacity headroom, not merely whether the current SLA happens to pass.
99. Your microservices application has 15 services. Under load, only the Order API becomes slow, but all other APIs appear healthy. How would you identify the root cause?
Scenario
Customer API → 200 ms
Catalog API → 150 ms
Inventory API → 180 ms
Payment API → 300 ms
Order API → 4 sec
Detailed Answer
I would trace the complete Order request.
Potential dependency chain:
Order API
↓
Inventory
↓
Payment
↓
Database
↓
Kafka
I would measure each segment.
For example:
Order API = 4 sec
Inventory call = 100 ms
Payment call = 3.2 sec
DB = 300 ms
Other = 400 ms
Now the likely bottleneck is Payment.
Distributed tracing
For microservices, distributed tracing is extremely valuable.
I want:
Trace ID
│
├── Order Service 4 sec
│
├── Inventory 100 ms
│
├── Payment 3.2 sec
│
└── Database 300 ms
I would also examine concurrency
Maybe Payment has:
Max connections = 100
while:
Order requests = 500 concurrent
This can create queueing.
Senior answer
Measure the critical path across service boundaries instead of blaming the service where the latency becomes visible.
100. Your load test generates 50,000 virtual users, but the application receives only 10,000 requests per second. The team believes the load tool is broken. What would you investigate?
Detailed Answer
I would first clarify:
Virtual users are not the same as requests per second.
Suppose each user performs:
Request
↓
Think time = 5 sec
↓
Request
↓
Think time
50,000 users can still generate relatively modest request rates.
Workload model
I would calculate:
Concurrency
+
Request rate
+
Response time
+
Think time
These are related but not identical.
A simplified relationship is:
Concurrency ≈ Throughput × Response/iteration time
I would investigate:
User behavior
Think time
Iterations
Requests per transaction
Connection reuse
Load-generator capacity
Network
Server-side request metrics
Example
If:
50,000 users
Average transaction = 10 sec
then the expected request rate depends heavily on how many requests each transaction generates.
Also verify server metrics
The load generator may report:
10,000 req/sec
while the application reports:
9,800 req/sec
That could be perfectly reasonable because of failed, cached, redirected, or filtered requests depending on architecture.
Senior answer
Performance engineers reason from workload characteristics, not from a single "virtual user" number.
101. Your team wants to run performance tests directly against production because the staging environment is much smaller. As Lead SDET, would you approve it?
Detailed Answer
I would not automatically approve or reject it.
I would first assess risk.
Production performance testing can affect:
Real users
Revenue
Database
External services
Infrastructure
Data
Preferred approach
Create a production-like environment:
Production
Architecture
↓
Staging / Performance
Environment
with comparable:
CPU
Memory
Database size
Network
Caching
Topology
Service dependencies
Configuration
If production testing is absolutely necessary
I would require controls.
For example:
Synthetic test accounts
Restricted endpoints
Limited traffic
Controlled time window
Monitoring
Abort thresholds
Rollback plan
Business approval
External-provider coordination
And ideally:
Production
↓
Small controlled load
↓
Monitor
↓
Increase gradually
↓
Stop automatically if thresholds exceeded
What I would never do
Run:
100,000 users
against production without a controlled plan simply because staging cannot reproduce production scale.
Senior answer
The goal is production-like performance testing, not production-risk performance testing.
102. Your load test shows that adding 4 application servers only improves throughput by 10%. The team expected nearly 4× improvement. What does this tell you?
Scenario
Before:
2 servers
→ 10,000 req/sec
After:
6 servers
→ 11,000 req/sec
Detailed Answer
This suggests the bottleneck is likely somewhere other than the application compute layer.
Potential bottlenecks:
Database
Cache
Message broker
Network
Load balancer
External API
Shared storage
Connection pool
Synchronization/locking
Example
Suppose:
Application servers:
2 → 6
DB:
CPU = 100%
The application servers can increase, but all of them compete for the same database.
┌── App 1 ──┐
├── App 2 ──┤
├── App 3 ──┤
├── App 4 ──┤
├── App 5 ──┤
└── App 6 ──┘
│
▼
Database
100% CPU
Adding more application nodes therefore produces diminishing returns.
I would investigate scalability efficiency
A useful question is:
2 servers → 10K
4 servers → ?
6 servers → 11K
8 servers → ?
This helps identify where scaling stops being effective.
I would also examine:
CPU utilization
DB utilization
Request queues
Lock contention
Network
Connection pools
External dependencies
Senior-level concept
This is a scalability bottleneck.
The system is not scaling linearly because another shared resource is limiting throughput.
Senior answer
If adding compute doesn't significantly increase throughput, look for a shared bottleneck or serialized part of the system.
103. You are asked to create a test strategy for a new banking application with only 6 weeks before release. How would you decide what to test?
Scenario
The application contains:
Login
Account Management
Fund Transfer
Beneficiary Management
Statements
Notifications
Admin Portal
You have:
6 weeks
5 SDETs
8 developers
1 QA environment
Management asks:
"Can you test everything before release?"
Detailed Answer
I would not start with the question :
"How many test cases can we execute?"
I would start with:
"What failures would cause the highest business/customer impact?"
I would create a risk-based test strategy.
Step 1 — Identify critical business journeys
For example:
High Risk
├── Login
├── Fund Transfer
├── Beneficiary Creation
└── Account Balance
Medium Risk
├── Statements
└── Notifications
Lower Risk
└── UI preferences
Step 2 — Assess risk
Risk can be considered using:
Risk = Probability × Impact
For example:
Feature Probability Impact Priority
Fund Transfer High Critical P0
Login Medium Critical P0
Statements Medium Medium P1
UI Preferences Low Low P3
Step 3 — Select test layers
Fund transfer might receive:
Unit
+
API
+
Integration
+
Database validation
+
UI E2E
+
Performance
+
Security
Whereas a minor UI preference may need only:
Component/UI
Step 4 — Define release gates
Example:
P0 tests → 100% PASS
P1 tests → ≥ 98% PASS
Critical defects → 0
Security blockers → 0
Performance SLA → PASS
Lead-level answer
I would communicate that 100% testing is impossible in six weeks, but 100% of critical risk areas can be systematically addressed.
The strategy should optimize risk reduction, not test-case count.
104. Your organization has 10,000 automated tests, but escaped production defects are increasing every quarter. Management asks, "Why aren't our 10,000 tests protecting us?" How would you answer?
Detailed Answer
I would challenge the assumption:
10,000 tests
≠
10,000 useful quality checks
I would investigate:
Coverage relevance
Test quality
Duplicate coverage
False positives
Flakiness
Production scenarios
Test-data realism
Missing integration paths
Missing negative scenarios
Example
Suppose:
10,000 tests
but:
3,000 → duplicate scenarios
2,000 → low-value UI checks
1,000 → flaky
2,000 → outdated
The actual meaningful coverage may be much smaller.
I would compare automation with production failures
Create a defect taxonomy:
Production defects
│
├── Functional
├── Integration
├── Data
├── Performance
├── Security
├── Configuration
└── Environment
Then ask:
Which categories are escaping our test strategy?
If 40% of escaped defects are integration failures, adding another 1,000 UI tests probably won't help.
Introduce a defect-to-test feedback loop
Production defect
↓
Root cause
↓
Why wasn't it detected?
↓
Missing test?
Wrong test layer?
Wrong environment?
Missing monitoring?
↓
Add preventive control
Senior answer
Automation volume is not a quality metric by itself. Coverage of meaningful risk and defect-prevention effectiveness matter much more.
105. A critical feature has 50 possible input combinations, but testing all combinations takes 3 days. How would you reduce the test effort without creating unacceptable risk?
Detailed Answer
I would use risk-based techniques and combinatorial testing rather than blindly testing every combination.
Suppose the feature depends on:
Browser
Country
User type
Payment method
Currency
Device
Testing every combination may create thousands of cases.
I would identify:
Critical combinations
Boundary values
Invalid combinations
Known high-risk combinations
Pairwise/multi-way interactions
Example
Instead of:
10 × 5 × 4 × 3 × 3 = 1,800 combinations
I might use a pairwise strategy to cover important interactions, supplemented with business-critical combinations.
But I would NOT blindly apply pairwise testing.
For financial functionality:
High-value transaction
+
International currency
+
Corporate user
+
Specific payment method
may require an explicit test even if a combinatorial algorithm doesn't prioritize it.
Strategy
Business-critical scenarios
+
Boundary scenarios
+
Negative scenarios
+
Pairwise coverage
+
Historical defect combinations
Senior answer
Combinatorial reduction should reduce redundant coverage, not remove business-critical risk.
106. Developers complain that SDET tests block their pull requests for 30–45 minutes. Product management wants faster releases. How would you redesign the quality gates?
Detailed Answer
I would analyze the pipeline rather than simply removing tests.
I would classify tests by:
Speed
Stability
Risk
Diagnostic value
Then create progressive gates.
PR gate
Unit
+
Component
+
Affected API
+
Critical smoke
Target:
< 10 minutes
Main branch
Broader integration
+
API regression
+
Selected E2E
Nightly
Full regression
+
Performance
+
Cross-browser
+
Extended scenarios
Release
Critical business flows
+
Security
+
Performance
+
Production smoke
Important
I would not remove a test simply because it is slow.
I would ask:
Can this validation happen at a cheaper layer?
For example:
UI test = 2 minutes
API equivalent = 3 seconds
Move the business-rule validation to API and retain only the critical UI journey.
Senior answer
Quality gates should be progressive: fast feedback early, deeper confidence later.
107. A developer says, "QA owns quality. Developers just need to write code." As Lead SDET, how would you change this mindset?
Detailed Answer
I would not solve this through confrontation.
I would establish shared quality ownership.
Quality should be considered throughout:
Requirement
↓
Design
↓
Development
↓
Testing
↓
Deployment
↓
Production
Shift-left
Before development starts:
Acceptance criteria
Testability
Observability
Failure scenarios
API contracts
Performance expectations
should be discussed.
Example
Instead of:
Developer builds feature → QA finds 20 defects
move toward:
Developer + SDET
↓
Risk analysis
↓
Test design
↓
Implementation
↓
Automated validation
Definition of Done
A feature isn't complete merely because:
Code compiled
It might require:
Unit tests
API tests
Automation
Observability
Security checks
Performance criteria
Documentation
depending on the feature.
Metrics
I would encourage:
Defect escape rate
Rework
First-pass CI success
Flaky tests
Mean time to detect
Mean time to recover
rather than:
Number of defects found by QA
Senior answer
QA/SDET should be a quality engineering function, not a final inspection department.
108. A product manager says, "We achieved 90% automation coverage, so we're ready for release." You believe the release is still high-risk. How would you explain why?
Detailed Answer
First, I would clarify what "90% automation coverage" actually means.
It could mean:
90% of test cases automated
which doesn't necessarily mean:
90% of business risk covered
Example
Suppose:
Authentication
→ 95% covered
Fund transfer
→ 60% covered
Rare but critical failure handling
→ 10% covered
Overall automation may still be 90%.
But the release risk remains high.
I would report coverage across dimensions:
Functional coverage
Business-risk coverage
Code coverage
API coverage
Integration coverage
Critical-path coverage
Negative-path coverage
Platform/browser coverage
Performance coverage
Security coverage
Example dashboard
Automation coverage 90%
Critical business paths 100%
High-risk scenarios 95%
API coverage 92%
Integration coverage 80%
Performance PASS
Security PASS
Known critical defects 0
That is much more meaningful.
Senior answer
Coverage is multidimensional. A single percentage can create a false sense of security.
109. Your application has thousands of test cases, but every release has only 2 hours available for regression. How would you determine which tests run during release?
Detailed Answer
I would implement risk-based regression selection.
Each test could have metadata such as:
Business criticality
Feature
Risk
Execution time
Failure history
Defect detection history
Dependencies
Last execution
Then create tiers.
Example
P0 — Release blockers
100 tests
P1 — High-risk regression
500 tests
P2 — Extended regression
1,500 tests
P3 — Full regression
Remaining tests
Release execution
P0
↓
P1
↓
Performance/security checks
↓
Release decision
Full regression can run:
Nightly
Weekend
Post-release
Dynamic selection
If today's change affects:
Payment Service
prioritize:
Payment tests
Order tests
Refund tests
Financial DB validations
Payment contracts
Critical checkout E2E
But maintain a safety net
A test-selection mechanism itself must be validated.
Otherwise you may accidentally exclude important tests forever.
Senior answer
Release regression should maximize risk coverage within the available time, while full regression remains part of the broader quality strategy.
110. Production has 99.99% availability, but customers still report that the checkout experience is unreliable. Management says, "Our availability is excellent." What would you investigate?
Detailed Answer
I would distinguish service availability from user-journey reliability.
A system may have:
API availability = 99.99%
but:
Checkout success rate = 97%
because checkout depends on multiple steps.
For example:
Login
↓
Cart
↓
Inventory
↓
Payment
↓
Order
↓
Confirmation
A failure in any step can break the user journey.
Google's SRE guidance emphasizes defining service indicators and objectives around behaviors that matter to users, rather than relying only on infrastructure-level metrics.
I would create journey-level indicators
For example:
Checkout success rate
Payment success rate
Order completion rate
Checkout latency
Cart-to-order conversion
Example
API uptime = 99.99%
Checkout success = 97%
Then investigate the missing 3%.
Potential causes:
Payment timeout
Inventory race
Frontend error
Session expiration
Third-party failure
Data inconsistency
Lead-level insight
The customer doesn't care that:
order-service-03
was available.
They care:
"Could I successfully complete my purchase?"
Senior answer
Quality engineering must measure critical user outcomes, not just component health.
111. Your organization releases every two weeks, but escaped defects have doubled despite increasing the QA team from 5 to 15 people. What would you investigate?
Detailed Answer
I would avoid assuming:
More QA
=
Fewer production defects
The problem may be systemic.
I would analyze:
Requirements
Architecture
Development practices
Code review
Test strategy
Automation
Environment
Deployment
Observability
Production feedback
Build a defect escape analysis
For every escaped defect:
Defect
↓
Root cause
↓
Why wasn't it detected?
↓
Where should it have been detected?
Example:
Production bug
↓
API contract changed
↓
No contract test
↓
Should have been caught in CI
Another:
Production bug
↓
Rare concurrency issue
↓
No load/concurrency testing
↓
Needs performance test
Another:
Production bug
↓
Bad deployment configuration
↓
Tests passed
↓
Needs deployment validation
Then prioritize systemic improvements
Contract testing
Test-data improvements
CI quality gates
Observability
Shift-left testing
Performance testing
Production monitoring
Senior answer
When escaped defects increase despite more testers, the problem is often the quality system—not the number of people executing tests.
112. You are asked to define the quality strategy for a product where teams deploy 50 times per day. Traditional full regression is impossible. What would you design?
Scenario
50 deployments/day
20 microservices
Multiple development teams
Continuous delivery
A traditional:
"Run regression before every release"
model doesn't scale.
Detailed Answer
I would design continuous quality validation.
Developer Commit
↓
Unit Tests
↓
Component Tests
↓
Contract Tests
↓
API Tests
↓
Security / Static Checks
↓
Build Artifact
↓
Deployment
↓
Smoke
↓
Canary
↓
Production Monitoring
Fast feedback
PR:
Minutes
Broader validation
Main branch:
Integration + selected E2E
Production confidence
Use:
Canary
Progressive rollout
Synthetic monitoring
Health checks
SLIs/SLOs
Automated rollback
SRE practices commonly use SLOs and error budgets to balance reliability with delivery velocity rather than requiring unrealistic 100% reliability.
Quality becomes continuous
Instead of:
Test
↓
Release
↓
Hope
you have:
Build
↓
Validate
↓
Deploy
↓
Observe
↓
Validate
↓
Expand rollout
↓
Observe
Example release gate
Canary 5%
↓
Error rate < 0.1%
p95 < 500 ms
Checkout success > 99%
↓
25%
↓
100%
If the SLO is breached:
Stop rollout
+
Rollback
+
Investigate
Senior answer
At high deployment frequency, quality cannot be a phase before release; it must become a continuous engineering control throughout the software lifecycle.
TQuestions 113–122
These are deliberately focused on Lead SDET responsibilities: test architecture, ephemeral environments, container failures, Kubernetes behavior, scalability, networking, observability, CI/CD, and production-like testing.
113. Your tests pass locally but fail intermittently when running inside Docker in CI. How would you investigate?
Scenario
Developer machine:
Tests → 100% PASS
CI:
Tests → 85–95% PASS
The failures are inconsistent.
Detailed Answer
I would first determine whether the problem is:
Application
Test
Container
Environment
Infrastructure
Timing
I would not immediately increase retries.
Step 1 — Compare environments
Compare:
Java/Node version
Browser version
OS
CPU
Memory
Environment variables
Timezone
Locale
Network
Dependencies
A common problem is:
Local:
8 CPU / 16 GB RAM
CI container:
1 CPU / 2 GB RAM
Tests may expose timing/resource problems.
Step 2 — Check container resources
Look for:
CPU throttling
Memory limits
OOM kills
Disk space
File descriptors
Process limits
Step 3 — Check test parallelism
Suppose locally:
4 workers
but CI:
20 workers
This may cause:
Resource contention
Port conflicts
Database contention
Test-data collision
Browser instability
Step 4 — Check dependencies
If containers start:
Application
Database
Redis
Kafka
the test may start before dependencies are actually ready.
This creates:
Container started
≠
Application ready
I would use proper readiness checks rather than arbitrary sleeps.
Step 5 — Collect diagnostics
For failures:
Container logs
Application logs
Test logs
Screenshots
Traces
Network information
Resource metrics
Senior answer
I would reproduce the CI runtime conditions locally and classify the failure before changing the test. Retries should hide transient infrastructure noise only when the underlying behavior is understood.
114. Your Kubernetes-based test environment randomly kills the application pod during a large test suite. What would you investigate?
Scenario
During testing:
Pod starts
↓
Tests run
↓
Pod restarts
↓
Tests fail
Kubernetes shows:
Restart Count: 3
Detailed Answer
My first step would be to determine why Kubernetes restarted the pod.
I would inspect:
Pod events
Container exit code
Previous container logs
Resource limits
Liveness probe
Readiness probe
Node events
Common causes
1. OOMKilled
Example:
Memory limit = 1 GB
Application uses = 1.3 GB
Kubernetes may terminate the container.
I'd investigate:
Memory leak
Heap configuration
Large test payloads
Caching
Concurrency
2. Liveness probe failure
Example:
Application becomes slow
↓
Liveness probe times out
↓
Kubernetes restarts pod
This can make the situation worse.
3. CPU throttling
If:
CPU request = 100m
CPU limit = 500m
but the application needs significantly more CPU, performance degradation may occur.
4. Node/resource pressure
The node may experience:
Memory pressure
Disk pressure
CPU pressure
Important distinction
I would separate:
Application failure
from:
Kubernetes health-management behavior
Senior answer
A pod restart is a symptom. The first job is to determine whether the restart was caused by resource exhaustion, health probes, application termination, or infrastructure pressure.
115. Your team creates a fresh test environment for every pull request. After a few months, CI becomes slow and cloud costs explode. How would you redesign the strategy?
Scenario
100 PRs/day
×
New Kubernetes environment
×
Multiple databases/services
Result:
High cost
Slow provisioning
Resource waste
Environment cleanup problems
Detailed Answer
Ephemeral environments are valuable, but they need lifecycle management.
I would analyze:
Provisioning time
Environment utilization
Environment lifetime
Test requirements
Infrastructure cost
Parallelism
Strategy
Use different environment types.
PR
↓
Lightweight ephemeral environment
↓
Affected-service testing
For broader integration:
Shared controlled environment
For release:
Production-like environment
Automatically destroy environments
For example:
PR opened
↓
Environment created
↓
Tests
↓
PR merged/closed
↓
Environment destroyed
Also add TTL protection:
Environment older than 24h
↓
Automatic cleanup
Reduce environment size
A PR may not need:
20 microservices
5 databases
3 replicas each
If only one service changed.
Use:
Real dependency
+
Mocked/non-critical dependency
where appropriate.
Senior answer
Ephemeral environments should be disposable, right-sized, observable, and automatically cleaned up.
116. Your Kubernetes deployment passes all automated tests, but users report intermittent 503 errors immediately after deployment. How would you investigate?
Scenario
Deployment:
Deployment → PASS
After release:
503 errors
Only during the first few minutes.
Detailed Answer
I would investigate the deployment/readiness path.
Potential sequence:
New pod created
↓
Pod receives traffic
↓
Application not fully initialized
↓
503
I would inspect
Readiness probe
Liveness probe
Startup probe
Service endpoints
Ingress
Load balancer
Pod startup time
Connection initialization
Cache warm-up
Database migration
Key distinction
A pod being:
Running
does not necessarily mean:
Ready to receive traffic
Example
Application requires:
20 seconds
to initialize.
But readiness check says:
HTTP 200
after only:
2 seconds
Traffic can arrive too early.
I would test rollout behavior
Old pods
↓
New pods
↓
Readiness
↓
Traffic shift
↓
Old pods termination
I would also verify:
RollingUpdate strategy
maxUnavailable
maxSurge
Production-style validation
Run:
Deployment
+
Continuous synthetic traffic
and monitor:
5xx
Latency
Availability
Pod readiness
Senior answer
Deployment testing must validate traffic readiness, not just whether Kubernetes reports the pod as running.
117. Your application works perfectly inside the Kubernetes cluster, but the API becomes unreachable from outside the cluster. What would you investigate?
Detailed Answer
I would trace the network path:
Client
↓
DNS
↓
Load Balancer / Ingress
↓
Service
↓
Pod
↓
Application
Check DNS
DNS resolution
TTL
Record
Hostname
Check ingress
Ingress rules
Host/path matching
TLS
Backend service
Annotations/configuration
Check Service
Service type
Port
TargetPort
Selector
Endpoints
A classic problem:
Service selector
↓
No matching pods
Then:
Service endpoints = empty
Check network policies
A Kubernetes NetworkPolicy may block traffic unexpectedly.
Check application binding
For example, application listens on:
127.0.0.1
instead of:
0.0.0.0
Then the service may not be able to reach it properly.
Senior debugging approach
I would test each hop independently:
Pod → localhost
Pod → Service
Pod → dependency
Outside → Load balancer
Outside → Ingress
This quickly narrows the failure domain.
Senior answer
Debug Kubernetes networking hop-by-hop instead of treating "API unreachable" as a single problem.
118. Your CI pipeline runs 500 Playwright tests in Kubernetes. Increasing workers from 5 to 30 makes the pipeline slower instead of faster. Why?
Detailed Answer
This is a classic parallelism saturation problem.
The assumption:
More workers = faster
is not always true.
Possible bottlenecks
CPU
Memory
Database
Network
Browser processes
Test environment
External APIs
CI runner
For example:
5 workers:
CPU = 60%
30 workers:
CPU = 100%
Memory = 95%
DB connections = exhausted
Now workers compete for resources.
I would measure
Worker count
Execution time
CPU
Memory
DB connections
Network
Browser startup time
Failure rate
Then find the optimal point.
Example:
Workers Time Failure Rate
5 40 min 1%
10 24 min 1%
15 18 min 2%
20 17 min 5%
30 22 min 12%
The optimum might be around:
15–20 workers
not 30.
Senior answer
Parallelism should be capacity-driven, not configured to the maximum possible worker count.
119. Your team wants every automated test to run against the same shared Kubernetes environment. After a while, tests randomly fail because one test changes data used by another. How would you solve this?
Detailed Answer
This is primarily a test isolation and environment contention problem.
Shared environments create:
Test A
↓
Changes data
↓
Test B
↓
Unexpected state
First principle
Tests should ideally be:
Independent
Repeatable
Isolated
Test-data isolation
Use unique identifiers:
user_<testId>
order_<testId>
rather than:
testuser
testorder
Parallel execution
Each worker can receive:
workerId
and generate isolated data.
Example:
worker-1 → customer_001
worker-2 → customer_002
worker-3 → customer_003
Environment isolation
For critical integration tests:
PR
↓
Ephemeral environment
or:
Namespace per test suite
where cost permits.
Database strategy
Depending on architecture:
Transaction rollback
Dedicated schema
Dedicated database
API-based cleanup
Data reset
Avoid blind cleanup
If Test A deletes:
customer_123
while Test B is using it, cleanup itself becomes a race condition.
Senior answer
Parallel automation requires both execution isolation and data isolation. Simply increasing Kubernetes resources won't solve shared-state problems.
120. Your application scales from 3 pods to 20 pods under load, but performance barely improves. How would you determine whether Kubernetes autoscaling is actually working?
Detailed Answer
I would separate two questions:
Did Kubernetes scale?
and:
Did scaling improve application capacity?
These are different.
First verify HPA behavior
Check:
Current replicas
Desired replicas
CPU utilization
Memory
Scaling metric
Scale-up events
Scale-down events
Then investigate why more pods don't improve throughput.
Possible causes:
Database bottleneck
External API bottleneck
Shared cache
Connection pool
Queue
Lock contention
CPU throttling
Network
Example
3 pods
→ 5K RPS
20 pods
→ 5.5K RPS
If:
DB CPU = 100%
then the application tier isn't the limiting factor.
Also verify autoscaling configuration
A poorly chosen metric can cause:
Slow scale-up
Over-scaling
Under-scaling
Oscillation
Important Lead SDET responsibility
I would test:
Scale-up time
Scale-down behavior
Maximum capacity
Recovery
Failure scenarios
Traffic spikes
Senior answer
Autoscaling should be validated as a system behavior: trigger → scale → capacity increase → stabilization → recovery.
121. Your team claims that "Docker makes tests reproducible," but the same container produces different results on different CI agents. How would you challenge that assumption?
Detailed Answer
Docker provides isolation, but it doesn't automatically guarantee complete determinism.
The container still depends on external factors.
For example:
Container
│
├── Host kernel
├── CPU
├── Memory
├── Network
├── Mounted volumes
├── Environment variables
├── External services
└── Time
I would compare CI agents
Docker version
Container image digest
CPU architecture
Kernel
Resources
Network
Environment variables
Mounted files
Secrets/config
Pin dependencies
Instead of:
latest
use immutable versions/digests where practical.
For example:
Browser version
Runtime version
Base image
Package versions
Check external dependencies
The same container may behave differently because:
Database state
External API
DNS
Network latency
Clock
differs between agents.
Check test randomness
If tests depend on:
Random data
Current time
Execution order
Thread scheduling
results may vary.
Use:
Deterministic seeds
Controlled clocks where appropriate
Isolated data
Stable dependency versions
Senior answer
Containerization improves reproducibility, but deterministic testing requires control of the container, host/runtime assumptions, dependencies, data, and timing.
122. You are asked to design a Kubernetes-based test infrastructure for a company with 50 microservices and 100 deployments per day. What would your architecture look like?
Scenario
Requirements:
50 microservices
100 deployments/day
500+ automated tests
Parallel execution
Fast feedback
Ephemeral environments
Production-like validation
Detailed Answer
I would design the platform around automation, isolation, scalability, and observability.
High-level architecture
Git / PR
│
▼
CI Pipeline
│
┌─────────┼─────────┐
▼ ▼ ▼
Unit API Contract
│ │ │
└─────────┼─────────┘
▼
Test Orchestrator
│
┌─────────┼─────────┐
▼ ▼ ▼
Namespace Namespace Namespace
│ │ │
Service A Service B Service C
│ │ │
└─────────┼─────────┘
▼
Integration
│
▼
E2E / Smoke
│
▼
Quality Decision
Ephemeral namespaces
For a PR:
PR #123
↓
Namespace: pr-123
↓
Deploy required services
↓
Run tests
↓
Collect artifacts
↓
Destroy namespace
Test orchestration
The orchestrator should support:
Parallel execution
Test sharding
Retry policy
Test selection
Resource limits
Timeouts
Artifact collection
Test data
Use:
Unique test identities
Factories
API setup
Controlled fixtures
Cleanup
Observability
Every test should have:
Build ID
PR ID
Test ID
Trace ID
Namespace
Pod
Service
Logs
Metrics
So a failed test can be traced:
Test failure
↓
Request
↓
Trace
↓
Service
↓
Pod
↓
Log
↓
Root cause
Environment lifecycle
Create
↓
Deploy
↓
Health check
↓
Test
↓
Collect evidence
↓
Destroy
with automatic cleanup for abandoned environments.
Scaling
The Kubernetes cluster itself should support:
Node autoscaling
Pod resource requests/limits
Test-worker scaling
Queue-based execution
Quality gates
PR
↓
Fast tests
↓
Contract/API
↓
Integration
↓
Critical E2E
↓
Deployment
↓
Smoke
↓
Canary
↓
Production monitoring
Lead SDET architecture principle
I would avoid creating one giant "QA Kubernetes cluster" where everything shares everything.
Instead:
Isolation
+
Repeatability
+
Scalability
+
Observability
+
Cost control
should drive the architecture.
Senior answer
The goal isn't merely to run tests inside Kubernetes. The goal is to build a scalable, disposable, observable quality platform that supports the organization's deployment velocity.
123. Two users can access the same API endpoint, but User A can retrieve User B's order by changing the orderId. How would you identify and automate this vulnerability?
Scenario
User A:
GET /api/orders/1001
→ 200 OK
User B owns:
orderId = 2001
User A changes the request:
GET /api/orders/2001
and receives User B's order.
Detailed Answer
This is a classic Broken Object Level Authorization (BOLA) scenario. OWASP identifies BOLA as API1:2023 and specifically emphasizes authorization checks whenever an API accesses an object using a user-supplied identifier.
O
OWASP Foundation
The important point is:
Authentication
≠
Authorization
User A is legitimately authenticated, but isn't authorized to access object 2001.
How I would test it
Create two users:
User A → Order A
User B → Order B
Then:
Authenticate User A
↓
Get Order A
↓
Replace orderId with Order B
↓
Send request
Expected:
403 Forbidden
or an appropriate non-disclosing response according to the API contract.
Automation design
I would build reusable authorization tests:
for each protected resource:
Owner → ALLOW
Other user → DENY
Admin → ALLOW
Unauthenticated → DENY
Important Lead-level consideration
Don't test only:
GET /orders/{id}
Look for all object references:
GET
PUT
PATCH
DELETE
Download
Export
Search/filter
Nested resources
For example:
DELETE /orders/2001
could be much more serious than merely reading the object.
Senior answer
I would create cross-user authorization tests systematically across object-oriented endpoints, not just test whether an authenticated user can access the endpoint.
124. Your application uses OAuth2/JWT authentication. Functional tests pass, but you're asked to verify that expired or invalid tokens cannot access protected APIs. How would you design the tests?
Detailed Answer
I would create a token-state matrix.
Token State Expected
Valid token Allow
Expired token Reject
Malformed token Reject
Missing token Reject
Wrong audience Reject
Wrong issuer Reject
Invalid signature Reject
Insufficient scope Reject
Revoked token, if supported Reject
Example
Valid token
↓
GET /api/account
↓
200
Then:
Expired token
↓
GET /api/account
↓
401
Important distinction
I would verify both:
Authentication
and:
Authorization
For example:
Valid token
+
Wrong scope
should not necessarily receive access.
JWT-specific validation
Depending on the architecture, I would verify expected validation of claims such as:
iss
aud
exp
nbf
iat
scope/roles
and the signature.
OWASP's testing guidance explicitly includes testing OAuth weaknesses and JWT/session-management behavior.
O
OWASP Foundation
Automation architecture
I would create token utilities:
TokenFactory
├── validToken()
├── expiredToken()
├── wrongAudienceToken()
├── insufficientScopeToken()
└── invalidToken()
Then tests can focus on behavior rather than token-generation details.
Senior answer
Security automation should verify the complete token lifecycle and authorization claims, not merely test that a valid JWT returns 200.
125. Your application has Admin, Manager, and Employee roles. A normal Employee discovers that an Admin API returns 200 OK. How would you investigate?
Scenario
Roles:
Admin
Manager
Employee
Endpoint:
POST /api/admin/users/{id}/disable
Employee calls it and receives:
200 OK
Detailed Answer
This could be Broken Function Level Authorization.
The user may be properly authenticated, but the API isn't enforcing the required role.
OWASP's API Security Top 10 explicitly identifies broken function-level authorization as a major API risk.
O
OWASP Foundation
I would create an authorization matrix
Endpoint Admin Manager Employee
View profile ✅ ✅ ✅
Create user ✅ ❌ ❌
Disable user ✅ Maybe ❌
View audit logs ✅ ✅ ❌
Change system settings ✅ ❌ ❌
Then automate the matrix.
Test pattern
for endpoint in protectedEndpoints:
for role in supportedRoles:
execute request
verify expected authorization
Important
I wouldn't rely only on UI visibility.
Even if the Employee UI doesn't display:
"Disable User"
the API must still reject direct requests.
Expected behavior
Typically:
Authenticated but unauthorized
→ 403
while:
Unauthenticated
→ 401
assuming those semantics are defined by the service.
Senior answer
Authorization must be enforced server-side. Hiding UI controls is not an authorization mechanism.
126. Your login system locks accounts after repeated failed passwords. A security team asks you to verify that the mechanism cannot be bypassed. What would you test?
Detailed Answer
I would test the complete authentication abuse-control behavior.
Basic scenario
Attempt 1 → Wrong
Attempt 2 → Wrong
Attempt 3 → Wrong
...
Threshold reached
→ Account locked
Then verify:
Correct password
→ Still blocked according to policy
until the documented unlock condition occurs.
I would test variations
Different IP
Different client
Different session
Different device
Case variations
Username normalization
Parallel login attempts
Password-reset flow
The goal is to determine whether the protection is enforced consistently.
Also test account enumeration
For example:
Existing user:
"Invalid password"
Non-existing user:
"Invalid password"
Ideally, authentication failures shouldn't unnecessarily reveal whether an account exists.
OWASP's current testing guide includes weak lockout, authentication bypass, account enumeration, and password-reset testing.
O
OWASP Foundation
Important Lead SDET boundary
I would perform these tests only in an authorized test environment with controlled accounts and agreed thresholds.
Senior answer
Authentication security tests should validate the complete abuse-control mechanism, including alternate authentication and recovery paths, not just the normal login attempt.
127. Your API response contains customer email, phone number, internal IDs, and other fields that the UI doesn't display. What would you investigate?
Scenario
API:
{
"id": 123,
"name": "John",
"email": "john@example.com",
"phone": "...",
"internalUserId": "...",
"adminNotes": "...",
"creditRiskScore": 87
}
The UI displays only:
name
Detailed Answer
I would investigate unnecessary data exposure and object-property authorization.
OWASP's API guidance specifically includes broken object property-level authorization, which addresses cases where APIs expose or allow modification of properties that the caller should not access.
O
OWASP Foundation
Test approach
First establish the data contract.
For each role:
Employee
Manager
Admin
define:
Allowed fields
Restricted fields
Writable fields
Read-only fields
Then automate response validation.
Example
Employee:
Allowed:
id
name
Should not receive:
creditRiskScore
adminNotes
internalUserId
Also test request manipulation
Suppose:
{
"name": "John",
"creditRiskScore": 100
}
If Employee submits that field, the server should not allow unauthorized modification.
Important distinction
There are two separate problems:
Excessive data returned
and:
Unauthorized property modification
Both need testing.
Senior answer
I would validate the API contract at the field level, not just verify HTTP status codes and a few business fields.
128. Your API accepts a url field for webhook configuration. A security engineer warns about SSRF. As an SDET, how would you test the feature safely?
Detailed Answer
This is an SSRF-related scenario. SSRF is explicitly included as API7:2023 in the OWASP API Security Top 10.
O
OWASP Foundation
The application receives:
POST /webhooks
{
"url": "..."
}
and the server later makes an outbound request.
Test strategy
I would use a controlled test endpoint that I own.
For example:
Application
↓
Controlled test server
↓
Capture request
I can verify:
Was a request made?
What method?
What headers?
What destination?
What response handling?
Security controls I would expect
Depending on requirements:
Allowlist destinations
URL validation
Scheme restrictions
Redirect controls
Network egress controls
DNS/IP validation
Authentication
Timeouts
Why redirects matter
An apparently safe URL may redirect somewhere unexpected.
Therefore I would test:
Allowed URL
Invalid URL
Unsupported scheme
Redirect
Unreachable destination
Timeout
Malformed URL
Also test resource exhaustion
A webhook endpoint shouldn't allow an attacker to consume unlimited resources.
OWASP's API guidance separately identifies unrestricted resource consumption as a major API risk.
O
OWASP Foundation
Senior answer
For SSRF testing, I would use controlled infrastructure and verify both application-level URL validation and network-level egress protections.
129. Developers accidentally committed an API key into the Git repository. The key was removed in the next commit. Management says, "The problem is fixed." Do you agree?
Detailed Answer
No.
Removing the secret from the latest file does not necessarily mean the secret is no longer present in repository history or other systems.
I would treat the exposed secret as compromised.
Immediate response
The correct sequence is generally:
Detect
↓
Revoke/rotate credential
↓
Investigate exposure
↓
Remove secret from repository/history as appropriate
↓
Audit usage
↓
Prevent recurrence
Test strategy
I would introduce automated secret detection in:
Pre-commit
Pull request
CI pipeline
Repository scanning
Container/image scanning
Better architecture
Secrets should come from:
Secret manager
rather than:
Source code
or:
Plain-text configuration
Also investigate
CI logs
Artifacts
Docker images
Build caches
Chat messages
Tickets
Deployment manifests
because secrets can leak through multiple channels.
Important SDET responsibility
I would add a CI security gate that fails when known secret patterns are detected, while avoiding noisy rules that developers routinely bypass.
Senior answer
Once a credential is exposed, removing the line of code doesn't make the credential safe. Rotate it first, then prevent recurrence.
130. A dependency used by your application receives a critical security vulnerability notification. The application tests are all green. Can the release proceed?
Detailed Answer
Not automatically.
Functional tests answer:
"Does the application behave correctly?"
They don't necessarily answer:
"Is the software supply chain safe?"
I would investigate
Affected dependency
↓
Version currently used
↓
Vulnerable version range
↓
Exploitability
↓
Whether vulnerable functionality is used
↓
Available fixed version
↓
Transitive dependencies
CI/CD strategy
I would introduce software composition analysis/dependency scanning.
Pipeline:
Build
↓
Dependency scan
↓
Policy evaluation
↓
Functional tests
↓
Security checks
↓
Release
Policy example
Critical exploitable vulnerability
→ Block release
High
→ Security review / conditional block
Medium
→ Track remediation
The exact policy should be determined by organizational risk and the vulnerability's actual applicability.
Container angle
If the application runs in containers, scan:
Application dependencies
+
Base image
+
OS packages
Important point
A security scanner finding isn't automatically equivalent to an exploitable production vulnerability.
I would expect triage to consider:
Severity
Exploitability
Exposure
Reachability
Business impact
Mitigation
Senior answer
Green functional tests do not override a critical supply-chain security risk. Security gates and functional gates answer different questions.
131. Your team wants to add security tests to every pull request, but the full security suite takes 3 hours. How would you integrate security testing without destroying developer feedback speed?
Detailed Answer
I would use a layered DevSecOps strategy.
Not every security test belongs in the PR gate.
PR
Fast checks:
Secret scanning
SAST
Dependency checks
Basic API security tests
Security unit tests
Target:
Minutes
Main branch
Broader:
API security regression
Container scanning
Expanded SAST
Dependency analysis
Nightly
Deeper testing:
DAST
Extended authorization tests
Security regression
Long-running scans
Release
Risk-based validation:
Critical security scenarios
Infrastructure configuration
Production-like DAST
Pen-test findings verification
OWASP's testing guidance covers authentication, authorization, session management, input validation, configuration, business logic, and API testing as distinct areas, supporting a layered approach rather than one giant security test suite.
Pipeline
PR
│
├── Secret scan
├── SAST
├── Dependency scan
└── Fast security tests
│
▼
Merge
│
├── Extended security
└── Integration tests
│
▼
Release
│
└── DAST / deeper validation
Senior answer
Security testing should be continuous and risk-based, with fast preventive controls early and deeper dynamic testing later in the pipeline.
132. A production incident occurs where a normal user accesses another customer's financial data. Functional tests had 95% pass rate. As Lead SDET, how would you determine what failed in the quality process?
Detailed Answer
I would treat this as a quality-system failure, not simply a missing test case.
Step 1 — Reproduce safely
Create:
User A
User B
Resource A
Resource B
Then reproduce the authorization boundary violation in a controlled environment.
Step 2 — Identify root cause
Questions:
Was authorization missing?
Was it implemented incorrectly?
Was only UI authorization tested?
Was the API tested?
Was cross-user access tested?
Was the endpoint newly introduced?
Was there a contract change?
Step 3 — Trace where the failure should have been caught
Possible layers:
Unit test
↓
Service test
↓
API test
↓
Integration test
↓
Security regression
↓
DAST
↓
Production monitoring
Step 4 — Add a permanent regression
For example:
User A
↓
Create Order A
User B
↓
Create Order B
User A
↓
Request Order B
↓
Expected: DENY
Step 5 — Expand the security matrix
Don't fix only one endpoint.
Search for:
GET /resource/{id}
PUT /resource/{id}
PATCH /resource/{id}
DELETE /resource/{id}
Download
Export
Nested resources
Step 6 — Improve the engineering process
Potential controls:
Authorization test templates
API security checklist
Security acceptance criteria
Threat modeling
Reusable authorization framework
CI security regression
Production detection
Most important Lead-level response
I would not say:
"QA missed a test."
I would ask:
"Why did our quality system allow an authorization boundary failure to reach production?"
That leads to a systemic improvement rather than blaming an individual.
133. Your production API's p95 latency suddenly increases from 400 ms to 2.5 seconds. There are no new application errors. How would you investigate?
Scenario
Before deployment:
p50 = 180 ms
p95 = 400 ms
p99 = 800 ms
After deployment:
p50 = 300 ms
p95 = 2.5 sec
p99 = 8 sec
Error rate:
0.05%
Detailed Answer
I would not assume:
"No errors means the system is healthy."
Latency itself is a production-quality signal.
I would first establish:
When did latency increase?
Which endpoint?
Which region?
Which instance?
Which customer segment?
Which request type?
Step 1 — Check deployment correlation
Deployment
↓
Latency increase?
If the timing matches, I would compare:
Old version
vs
New version
Step 2 — Check infrastructure
CPU
Memory
GC
Thread pools
Connection pools
Network
Step 3 — Check dependencies
Database
Redis
Kafka
External APIs
Step 4 — Distributed tracing
Suppose the trace shows:
API 2.5 sec
├── Inventory 100 ms
├── Payment 150 ms
├── Database 2.1 sec
└── Other 150 ms
The database becomes the primary investigation area.
Step 5 — Compare query performance
Check:
Query plans
Indexes
Locks
Connection pool
DB CPU
I/O
Step 6 — Reproduce
Run a controlled performance test against the same version/configuration.
Lead-level conclusion
I would create a timeline:
14:00 Deployment
14:05 p95 = 450ms
14:15 p95 = 900ms
14:30 p95 = 2.5sec
Then correlate it with system metrics.
Senior answer
Latency degradation without errors is still an incident. I would correlate deployment, latency, infrastructure, dependencies, and traces to identify where the additional latency is being introduced.
134. Your monitoring dashboard shows an API error rate of only 0.5%, but the business team reports that 8% of checkout attempts are failing. How is that possible?
Detailed Answer
This is a classic example of technical health vs business health.
The API may be returning:
HTTP 200
while the business operation actually fails.
For example:
Checkout request
↓
200 OK
↓
Payment declined
↓
Order not created
The infrastructure dashboard sees:
HTTP errors = 0
but the customer sees:
Checkout failed
I would introduce business-level SLIs
Examples:
Checkout success rate
Payment success rate
Order creation success rate
Registration completion rate
For example:
Checkout Success Rate =
Successful Orders / Checkout Attempts
Then correlate
Checkout failures
↓
Payment service
↓
Payment decline
↓
External provider
Important point
A technical metric such as:
HTTP 5xx
doesn't necessarily represent the complete business failure rate.
Senior answer
A system can be technically available while the business transaction is unavailable. Lead SDETs should monitor both technical and business-level indicators.
135. Your synthetic Playwright test fails once every 30 runs in production. Developers say it is "just a flaky test." How would you determine whether it's actually a production problem?
Scenario
Synthetic test:
Login
→ Search
→ Add item
→ Checkout
Failure rate:
~3%
Detailed Answer
I would not immediately classify it as test flakiness.
The test itself is a production monitoring signal.
Step 1 — Correlate failures
For every failure, collect:
Timestamp
Region
Browser
Build/version
Request IDs
Trace IDs
Screenshot
Video
Console logs
Network logs
Step 2 — Compare with production telemetry
Suppose synthetic failures occur when:
Checkout API p95 > 2 sec
That strongly suggests the test is detecting a real production condition.
Step 3 — Determine failure pattern
Only one region?
Only one browser?
Only after deployments?
Only during peak traffic?
Only one endpoint?
Example
Failures:
Asia region → 0.2%
US region → 0.1%
Europe → 12%
Now infrastructure/region-specific investigation becomes important.
If it really is test flakiness
You might discover:
Selector instability
Timing issue
Third-party UI
Test-data collision
But you should prove that.
Senior answer
A production synthetic test should be treated as an observability signal. Never label an intermittent failure "flaky" until you have correlated it with production telemetry.
136. One microservice reports 99.99% availability, but the customer's end-to-end checkout journey is failing. How would you investigate?
Scenario
Order Service → Healthy
Payment Service → Healthy
Inventory → Healthy
But:
Checkout Success Rate = 92%
Detailed Answer
I would investigate the entire customer journey.
Customer
↓
Frontend
↓
Cart
↓
Order
↓
Inventory
↓
Payment
↓
Order confirmation
Each service can individually look healthy while the workflow fails.
Example
Order API = 99.99%
Payment API = 99.99%
Inventory API = 99.99%
But:
Order created
↓
Inventory reservation
↓
Payment succeeds
↓
Order confirmation event lost
The individual APIs may still show healthy availability.
I would use distributed tracing
For failed transactions:
Trace
├── Frontend 100ms
├── Order 200ms
├── Inventory 150ms
├── Payment 300ms
└── Event publish FAILED
Also investigate asynchronous components
Kafka
Queues
Consumers
Dead-letter queues
Retry queues
Senior answer
Component availability does not guarantee workflow reliability. Observability must follow critical business journeys across synchronous and asynchronous boundaries.
137. Production logs contain thousands of errors, but engineers cannot determine which customer request caused each error. As Lead SDET, what would you recommend?
Detailed Answer
The logging strategy lacks request correlation.
Every request should have a correlation identifier.
For example:
X-Correlation-ID
or a platform-standard trace/request ID.
Desired flow
Customer Request
│
▼
API Gateway
│
Trace ID
│
├── Service A
│
├── Service B
│
├── Database
│
└── Kafka
Then an incident can be traced using:
traceId = abc123
Logs should contain useful context
For example:
timestamp
service
version
environment
traceId
requestId
endpoint
status
latency
error type
But avoid logging sensitive information such as:
Passwords
Tokens
Secrets
Full payment information
Sensitive personal data
SDET responsibility
I would add automated validation for logging requirements.
For example:
Every critical API request
→ correlation ID present
→ response contains/propagates ID
→ downstream calls preserve trace context
Senior answer
Observability isn't just about collecting more logs. The logs must allow engineers to connect an individual business transaction across distributed components.
138. Your production services have timestamps that differ by several seconds, making incident investigation difficult. How would you solve it?
Detailed Answer
I would investigate clock synchronization and timestamp standards.
Distributed systems depend heavily on consistent time representation.
Standardize
Use:
UTC
+
ISO 8601
+
Consistent timestamp format
Infrastructure
Ensure hosts/nodes synchronize time using approved time synchronization mechanisms.
Application
Use server-generated timestamps where appropriate and ensure logs contain:
Timestamp
Timezone/UTC
Trace ID
Service
Version
Important distinction
There are multiple timestamps:
Client time
API gateway time
Service time
Database time
Queue time
I would identify which timestamp is being used for each purpose.
Distributed tracing
Trace systems can provide more reliable ordering/context across services than trying to manually compare unrelated log files.
SDET validation
I could add a test that verifies:
Request
↓
Service A
↓
Service B
↓
Service C
and confirms trace/correlation context is propagated correctly.
Senior answer
Consistent time representation plus distributed tracing is essential for reconstructing events across a distributed system.
139. Your company wants SDETs to own synthetic production monitoring. How would you design it?
Scenario
Critical customer journey:
Login
→ Search
→ Add to cart
→ Checkout
Detailed Answer
I would create safe synthetic users and controlled data.
Architecture
Scheduler
↓
Synthetic Test Runner
↓
Production
↓
Application
↓
Metrics / Logs / Traces
↓
Alerting
Test design
Tests should be:
Short
Stable
Representative
Safe
Idempotent where possible
Example
Every 5 minutes:
Login
↓
Search known product
↓
Add synthetic item
↓
Create controlled transaction
For financial systems, I'd avoid real financial side effects and use approved non-production/sandbox mechanisms where possible.
Capture
Duration
Status
Step failure
Region
Browser
Version
Trace ID
Alerting
Don't alert on a single transient failure.
Example:
1 failure → record
2 failures → investigate
3 consecutive failures → alert
The exact threshold should be based on service criticality and noise characteristics.
Important
Synthetic monitoring shouldn't become a substitute for real-user monitoring.
Use both where appropriate:
Synthetic
+
Real User / Business Metrics
+
Infrastructure Metrics
Senior answer
Synthetic monitoring should continuously validate critical customer journeys while minimizing production side effects and alert noise.
140. Your company has defined an SLO of 99.9% successful checkout transactions. How would you use that SLO as a Lead SDET?
Detailed Answer
First, define the SLI clearly.
For example:
Successful Checkout Transactions
--------------------------------
Total Checkout Attempts
Suppose:
SLO = 99.9%
This allows:
0.1%
of transactions to fail within the defined measurement window.
Then connect testing to the SLO.
Pre-production
Performance tests validate:
Expected traffic
Latency
Error rate
Capacity
CI/CD
Critical quality gates may include:
Checkout regression
API tests
Contract tests
Performance checks
Production
Monitor:
Checkout success
Latency
Errors
Error budget
The SLO creates an error budget.
Conceptually:
SLO = 99.9%
↓
Allowed unreliability
↓
Error budget
If the service is consuming the budget rapidly, I would recommend greater release caution.
Google's SRE guidance describes SLOs as a way to define acceptable reliability and use error budgets to balance reliability against development velocity.
Lead SDET contribution
I would help establish:
Quality gates
Synthetic monitoring
Regression priorities
Performance thresholds
Release risk criteria
Senior answer
An SLO should influence both what we test before release and what we monitor after release. It becomes a measurable definition of acceptable quality.
141. A release passes every test in staging, but production latency increases by 300%. As Lead SDET, what would you change in the quality strategy?
Detailed Answer
I would first perform a production-vs-staging gap analysis.
Compare:
Traffic
Data volume
Database size
Infrastructure
Caching
Network
External dependencies
Configuration
Concurrency
Autoscaling
Example
Staging:
10K records
2 instances
100 users
Production:
500M records
20 instances
20K concurrent users
Passing staging doesn't prove production scalability.
Improvements
1. Production-like performance environment
Increase:
Data volume
Traffic
Infrastructure similarity
2. Production performance testing
Use controlled production testing where appropriate and approved.
3. Canary deployment
1–5%
↓
Observe
↓
25%
↓
Observe
↓
100%
4. Automated production gates
Monitor:
p95
p99
Error rate
Business success rate
CPU
DB latency
5. Add the incident to regression strategy
If production failure occurred because of:
Large DB volume
then future performance testing should include realistic data volume.
Senior answer
A production escape should change the test strategy, environment model, and release controls—not simply result in one additional regression test.
142. During a production incident, developers say the application is healthy, the database team says the database is healthy, and the SDET synthetic test says checkout is failing. How would you lead the investigation?
Detailed Answer
This is where a Lead SDET should act as a system-level investigator, not argue about which team is correct.
I would start with the customer transaction.
Customer
↓
Frontend
↓
Gateway
↓
Order
↓
Inventory
↓
Payment
↓
Database
↓
Events
Step 1 — Establish the exact failure
Timestamp
Region
User
Journey step
HTTP status
Business response
Trace ID
Step 2 — Follow the trace
For example:
Checkout
↓
Order Service 200 ms
↓
Inventory 150 ms
↓
Payment 250 ms
↓
Kafka Publish 20 ms
↓
Consumer 12 sec
Now the issue may be downstream asynchronous processing.
Step 3 — Compare signals
Application health → PASS
Database health → PASS
Synthetic checkout → FAIL
These aren't contradictory.
They measure different things.
Step 4 — Check business metrics
Checkout success ↓
Payment success ↓
Queue latency ↑
Now the evidence points toward a particular component.
Step 5 — Incident containment
Depending on impact:
Pause rollout
Rollback
Disable problematic feature
Route traffic
Increase capacity
Step 6 — After recovery
Perform:
Root-cause analysis
↓
Missing detection?
↓
Missing test?
↓
Missing metric?
↓
Missing alert?
↓
Missing deployment gate?
Then permanently improve the system.
Lead-level answer
During an incident, I would use the synthetic test as one signal and correlate it with traces, metrics, logs, infrastructure, and business indicators. The objective is to establish the failure path, not prove which team is responsible.
Questions 143–152 will cover:
Parallel tests modify the same customer/order data — how would you redesign test-data isolation?
Production-like database has billions of records, but staging has only thousands — how would you create realistic test data?
A test passes with freshly created data but fails with old production-like data — how would you investigate?
Multiple teams share one QA environment and deployments frequently break other teams' tests — how would you design the environment strategy?
A test creates data through five microservices and cleanup occasionally fails — how would you make test setup/cleanup reliable?
GDPR/PII requirements prevent copying production data to lower environments — how would you create realistic but safe test data?
Tests pass individually but fail when executed in parallel because of database state — how would you solve it?
A database migration works on an empty database but fails on an existing production-like database — how would you test it?
Different test environments have different service versions/configurations — how would you detect and prevent environment drift?
Design a scalable test-data and environment strategy for 50 microservices and hundreds of daily deployments — Lead SDET architecture question.
The core Lead SDET mindset
Instead of:
Test
↓
Create some data
↓
Run test
↓
Delete data
we'll look at:
Test Data Platform
│
┌───────────────┼───────────────┐
▼ ▼ ▼
Data Factory Data Isolation Data Masking
│ │ │
└───────────────┼───────────────┘
▼
Test Environment
│
┌──────────┼──────────┐
▼ ▼ ▼
API Tests UI Tests E2E Tests
│ │ │
└──────────┼──────────┘
▼
Cleanup / TTL
Absolutely. Let's continue.
Category 13 — Test Data Management & Environment Strategy
10 Real / Scenario-Based Questions
Questions 143–152
These are focused on Senior SDET / Lead SDET decision-making rather than basic test-data concepts.
143. Your 200 parallel automated tests are using the same customer account, and tests randomly fail because one test changes the account state. How would you redesign the test-data strategy?
Scenario
Current setup:
200 tests
↓
customer = "testuser"
↓
Parallel execution
↓
Random failures
One test might:
id="9x4b6f"
Change address
while another expects:
id="s6j2k1"
Original address
Detailed Answer
The fundamental problem is shared mutable state.
I would not solve this by adding retries.
Preferred approach — Test Data Factory
Create data specifically for each test:
Test
↓
Data Factory
↓
Unique Customer
↓
Execute Test
↓
Cleanup / TTL
For example:
worker-01 → customer_test_001
worker-02 → customer_test_002
worker-03 → customer_test_003
Better than random strings
Use deterministic identifiers:
testRunId
workerId
testId
Example:
customer_<runId>_<workerId>_<testId>
This makes failures easier to reproduce.
Data ownership
Each test should know:
What data it created
What data it owns
How to clean it
Alternative approaches
Depending on the system:
Dedicated database
Dedicated schema
Namespace
Transaction rollback
API-level cleanup
Database fixture
Important Lead-level point
I would classify tests based on data isolation requirements:
Fully isolated
Partially shared
Read-only shared
Environment-global
Shared data should be minimized.
Senior answer
Parallel test execution requires data isolation. Every test should own its mutable data whenever practical, with deterministic creation and reliable cleanup.
144. Production contains 500 million customer records, but your QA environment has only 100,000. Performance tests pass in QA but fail badly in production. How would you address this?
Detailed Answer
The problem isn't necessarily the application.
It may be data-volume mismatch.
For example:
QA:
100K records
Production:
500M records
A query that performs well on 100K records may behave very differently at production scale.
I would identify important data characteristics
Not just total row count.
For example:
Total records
Data distribution
Hot records
Old records
Large records
Indexes
Cardinality
Relationships
Partitioning
Create representative datasets
Instead of copying production blindly, create synthetic data with similar characteristics:
500M logical records
+
Realistic distribution
+
Realistic relationships
+
Realistic index/cardinality behavior
Important
A dataset of:
500M identical records
may not reproduce production behavior.
You need realistic distribution.
Test multiple scales
For example:
100K
1M
10M
100M
500M
Then measure:
Latency
Throughput
DB CPU
Memory
IO
Query plans
Senior answer
Performance test data must represent production scale and data distribution, not merely contain a similar number of rows.
145. A test passes when it creates a new customer, but fails when executed against a customer that has existed for three years. How would you investigate?
Scenario
Fresh customer:
Created today
→ PASS
Production-like customer:
Created 3 years ago
→ FAIL
Detailed Answer
I would suspect state-dependent behavior.
Possible differences:
Account history
Status
Legacy fields
Old schema
Subscriptions
Transactions
Permissions
Preferences
Archived records
Migration state
Compare the two records
I'd build a diff:
New customer
vs
3-year-old customer
Compare:
Schema
Values
Relationships
Metadata
Status
Audit history
Investigate business rules
Maybe:
Customer age > 2 years
activates a different workflow.
Or:
Legacy customer
has a field that newer records don't.
Database investigation
Check:
Migration history
Null/default values
Indexes
Constraints
Triggers
Stored procedures
Test strategy
Create test personas:
New customer
Legacy customer
Inactive customer
High-value customer
Customer with large transaction history
Customer with migrated data
This gives better coverage than always creating fresh entities.
Senior answer
Real systems have historical state. Test data should model lifecycle states, not just newly created records.
146. Five teams share one QA environment. Team A deploys a new service version and suddenly Team B's tests fail. How would you redesign the environment strategy?
Detailed Answer
The fundamental issue is environment contention.
A single shared environment creates:
Deployment coupling
Data coupling
Configuration coupling
Test instability
I would introduce environment tiers
Developer
↓
PR / Ephemeral
↓
Team Integration
↓
Shared Integration
↓
Pre-production
↓
Production
PR environments
For isolated changes:
PR #123
↓
Namespace/environment
↓
Tests
↓
Destroy
Team environments
Critical teams may have:
Team A → QA-A
Team B → QA-B
depending on infrastructure cost.
Shared environment
Use it primarily for:
Cross-service integration
System-level validation
Release candidates
rather than every developer test.
Environment ownership
Each environment should have:
Owner
Purpose
Supported versions
Deployment rules
Data policy
Reset strategy
TTL
Senior answer
A shared QA environment should be treated as a managed platform, not as an unlimited sandbox. Isolation should increase as test criticality and deployment frequency increase.
147. Your end-to-end test requires data to be created across five microservices. Sometimes the test data is created successfully in four services but fails in the fifth. How would you design reliable setup and cleanup?
Scenario
Customer
↓
Account
↓
Subscription
↓
Payment profile
↓
Order
Setup fails at:
Payment profile
Now partial data remains.
Detailed Answer
I would avoid putting all setup logic directly into UI tests.
Instead, create a test data orchestration layer.
Test
↓
Test Data Service
↓
Service APIs
↓
Required entities
Use dependency-aware setup
Customer
↓
Account
↓
Subscription
↓
Payment
↓
Order
If step 4 fails, cleanup should understand what was already created.
Track created resources
Example:
created:
customer = C100
account = A100
subscription = S100
If payment creation fails:
cleanup:
S100
A100
C100
Idempotent cleanup
Cleanup should tolerate:
Already deleted
Partial creation
Retries
Timeouts
TTL-based safety net
Even if cleanup fails:
Created:
2026-08-12 15:00
automatically expire after:
24 hours
or an appropriate period.
Important
Don't make cleanup itself dependent on the same broken service whenever possible.
Senior answer
Test setup should be orchestrated, dependency-aware, traceable, and cleanup-safe. A failed test must not leave uncontrolled state behind.
148. Your organization prohibits copying production PII into QA. However, QA needs realistic customer data. What would you implement?
Detailed Answer
I would use synthetic data or properly de-identified/masked data, depending on requirements.
Preferred option
Generate synthetic data:
Customer
Address
Phone
Orders
Payments
Transactions
with realistic relationships.
Example:
Customer
↓
10 Orders
↓
50 Order Items
↓
Payment History
If production-derived data is necessary
Apply approved de-identification/masking controls.
For example:
Real email
→ synthetic email
Real phone
→ synthetic phone
Real name
→ synthetic name
But simple masking isn't always sufficient.
Re-identification risk
Suppose you change:
Name
but leave:
Rare address
DOB
Transaction pattern
the individual may still be identifiable.
Therefore
I would involve:
Security
Privacy
Compliance
Data governance
before implementing production-data extraction.
Validate the resulting dataset
Ensure:
Referential integrity
Distribution
Relationships
Edge cases
Required states
are preserved.
Senior answer
The objective is not simply to hide names. The objective is to preserve test usefulness while eliminating unacceptable privacy and re-identification risk.
149. Tests pass individually but fail when executed in parallel because database state is shared. How would you diagnose and solve it?
Detailed Answer
First I would prove that parallelism causes the issue.
Run:
Sequential
→ PASS
then:
Parallel
→ FAIL
Find the shared resource
Possible shared state:
Database row
Schema
Sequence
Cache
File
Queue
User
Configuration
Add test identifiers
Every test execution gets:
runId
workerId
testId
Then trace database records back to their owner.
Example
Instead of:
UPDATE customer
SET status='ACTIVE'
WHERE id=100;
use a test-owned customer:
customer_<runId>_<testId>
Database isolation options
Depending on architecture:
Transaction rollback
Dedicated schema
Database clone
Test container
Namespace
Per-worker dataset
Watch for hidden shared state
Even if customer records are isolated, this may still be shared:
SELECT MAX(order_id)
or:
UPDATE global_configuration
Race condition example
Test A → reads value 10
Test B → reads value 10
Test A → writes 11
Test B → writes 11
Expected:
12
Actual:
11
That's a concurrency problem, not simply a test problem.
Senior answer
I would identify every mutable shared resource and establish ownership or isolation. Parallelism exposes hidden coupling that sequential execution can conceal.
150. A database migration works on a new empty database but fails on a production-like database containing years of data. How would you test the migration?
Detailed Answer
Testing only:
Empty DB → Migration
is insufficient.
I would create multiple database states.
Test matrix
Empty DB
Fresh DB
Current QA DB
Production-like DB
Large DB
Legacy DB
Partially migrated DB
Validate
Before:
Schema
Record counts
Constraints
Indexes
Relationships
After:
Schema
Data integrity
Record counts
Indexes
Constraints
Application compatibility
Important checks
Null handling
Duplicate data
Existing invalid records
Large tables
Long-running operations
Locking
Downtime
Rollback
Backward compatibility
For rolling deployments:
Old application
↕
New database schema
↕
New application
The old and new application versions may temporarily coexist.
Example
Instead of immediately removing:
old_column
use:
Add new column
↓
Deploy compatible application
↓
Backfill
↓
Validate
↓
Switch reads/writes
↓
Remove old column later
Senior answer
Database migration testing must cover realistic existing data, compatibility during rollout, performance, integrity, and rollback—not just whether the migration succeeds on an empty database.
151. Your staging environment works correctly, but its configuration differs from production in 30 places. How would you detect and prevent environment drift?
Detailed Answer
I would treat configuration as version-controlled infrastructure, not tribal knowledge.
Create a configuration inventory
Application version
Environment variables
Feature flags
Database settings
External endpoints
Resource limits
Security settings
Network rules
Secrets references
Infrastructure as Code
Where practical:
Environment definition
↓
Git
↓
CI/CD
↓
Environment
This allows differences to be reviewed.
Configuration diff
Automatically compare:
Staging
vs
Production
and classify differences:
Expected
Allowed
Unexpected
Critical
Example
Expected:
database.hostname
must differ.
Unexpected:
feature.enableNewCheckout
should be identical unless intentionally different.
Policy checks
CI can fail if:
Required configuration missing
Unexpected configuration added
Unsafe value detected
Resource limit absent
Important
Don't blindly make staging identical to production.
Some differences are necessary:
Credentials
URLs
Capacity
Third-party sandbox endpoints
The goal is:
Controlled difference, not accidental difference.
Senior answer
Environment parity should mean intentional, version-controlled differences—not that every environment is literally identical.
152. You have 50 microservices, 500 automated tests, and hundreds of deployments per day. Design a scalable test-data and environment strategy.
Detailed Answer
This is the Lead SDET architecture question.
I would design around:
Isolation
Automation
Scalability
Repeatability
Observability
Cost
Security
High-level architecture
Git / PR
│
▼
CI/CD
│
┌─────────┼─────────┐
▼ ▼ ▼
Unit API Contract
│ │ │
└─────────┼─────────┘
▼
Environment Manager
│
┌─────────┼─────────┐
▼ ▼ ▼
PR-101 PR-102 PR-103
Env Env Env
│ │ │
▼ ▼ ▼
Test Data Test Data Test Data
Factory Factory Factory
│ │ │
└─────────┼─────────┘
▼
Integration/E2E
│
▼
Observability
Test-data platform
Provide reusable factories:
CustomerFactory
OrderFactory
PaymentFactory
SubscriptionFactory
ProductFactory
Example:
customer = customerFactory.create({
status: "ACTIVE",
type: "PREMIUM"
})
Unique ownership
Each test gets:
runId
workerId
testId
and therefore unique data.
Environment strategy
PR
Ephemeral environment
Only deploy required services where architecture permits.
Integration
Shared controlled environment
Release
Production-like environment
Data lifecycle
Create
↓
Use
↓
Validate
↓
Cleanup
↓
TTL safety net
Production data
Never casually copy production PII.
Use:
Synthetic data
+
Approved masked/de-identified datasets
Environment drift
Use:
Infrastructure as Code
+
Configuration as Code
+
Automated validation
Scalability
Support:
Parallel test execution
Test sharding
Dynamic environment creation
Resource quotas
Cluster autoscaling
Observability
Every test should be traceable to:
Build
↓
Environment
↓
Test
↓
Data
↓
Service
↓
Trace
↓
Logs
Cost management
Ephemeral environments need:
TTL
Auto-cleanup
Right-sized resources
Idle detection
Namespace quotas
Quality architecture
I would divide tests by purpose:
PR
├── Unit
├── Contract
├── Fast API
└── Critical regression
Post-merge
├── Integration
├── Broader API
└── E2E
Release
├── Full regression
├── Performance
└── Production-like validation
Lead-level answer
At 50 microservices and hundreds of deployments, the solution cannot be "create more test scripts." We need a test-data platform and environment platform that are automated, isolated, observable, scalable, and cost-controlled.
🔥 The Lead SDET principle
A weak answer is:
"I'll create test data before every test."
A stronger answer is:
Test
↓
Identify data requirements
↓
Generate isolated data
↓
Execute
↓
Validate
↓
Trace ownership
↓
Cleanup
↓
TTL safety net
A Lead SDET thinks one level higher:
Test Data Platform
+
Environment Platform
+
CI/CD Platform
+
Observability
↓
Scalable Quality Engineering
153. A consumer expects customerName in an API response. The provider removes the field because the provider team says, "The UI doesn't use it anymore." How would you prevent this breaking change?
Scenario
Current response:
{
"customerId": "C101",
"customerName": "John",
"status": "ACTIVE"
}
Provider changes it to:
{
"customerId": "C101",
"status": "ACTIVE"
}
Provider tests pass.
Consumer starts failing.
Detailed Answer
The provider team cannot determine compatibility solely from its own tests.
The important question is:
Does any consumer depend on this field?
This is exactly where consumer-driven contract testing helps.
In a consumer-driven model:
Consumer
↓
Defines interaction it actually needs
↓
Contract
↓
Provider verifies contract
Pact describes this model as having the consumer specify expected interactions and the provider verifying those interactions.
I would implement
Consumer contract
↓
Published to contract repository/broker
↓
Provider CI
↓
Contract verification
↓
PASS / FAIL
If customerName is part of the consumer's required interaction, removing it should cause provider verification to fail.
But there is an important nuance
I would not create a contract that asserts every field returned by the provider.
For example, this can become brittle:
{
"customerId": "C101",
"customerName": "John",
"status": "ACTIVE",
"createdAt": "...",
"internalCode": "...",
"serverTime": "..."
}
If the consumer doesn't use those fields, the contract shouldn't unnecessarily constrain them.
Pact's guidance explicitly recommends keeping consumer contracts as loose as possible while still protecting the consumer's real expectations.
P
Pact Docs
Senior answer
I would make the consumer's actual dependency explicit through a contract and verify that contract against the provider before deployment. I would not make the contract stricter than the consumer's real needs.
154. Your contract tests are passing, but an end-to-end test fails in production after a provider deployment. How can that happen?
Detailed Answer
This is a very important Lead SDET question.
A contract test doesn't prove that the entire distributed workflow works.
It proves a particular interaction between consumer and provider.
For example:
Consumer
↓
Provider
may be contract-compatible.
But production is:
Frontend
↓
API Gateway
↓
Order Service
↓
Inventory
↓
Payment
↓
Kafka
↓
Notification
A contract test might validate:
Order → Payment
while the production failure is:
Payment → Kafka → Notification
Other possible gaps
Business logic
Contract:
200 + expected fields
but business workflow:
Payment successful
BUT
Order not finalized
Authentication/authorization
Contract may validate the response structure but not the complete production identity/permission setup.
Configuration
Contract environment:
Provider URL = correct
Production:
Wrong routing/configuration
Data
Contract test:
Customer = simple fixture
Production:
Legacy customer
Complex account
Asynchronous behavior
Contract tests may not prove:
Event published
→ consumed
→ processed
→ eventually consistent state
Therefore
I would maintain layers:
Unit
↓
Contract
↓
API / Integration
↓
E2E
↓
Synthetic production monitoring
Senior answer
Contract testing reduces integration risk; it doesn't prove the entire system is functionally correct. I would investigate the exact boundary that the contract did not cover rather than conclude that contract testing failed.
155. You have 50 microservices, and every team deploys independently. How would you introduce contract testing without creating a huge maintenance burden?
Detailed Answer
I would avoid creating contracts for every possible API behavior.
Instead:
Consumer
↓
Document actual dependency
↓
Create focused contract
↓
Provider verifies it
Pact's guidance specifically distinguishes contract testing from broad provider functional testing; contracts should focus on consumer/provider expectations rather than becoming another full functional test suite.
P
Pact Docs
Architecture
Service A
│
├── Consumer contract → Service B
│
├── Consumer contract → Service C
│
└── Consumer contract → Service D
Service B
├── verifies A
├── verifies E
└── verifies F
Central contract repository
Use a broker/repository to store:
Consumer
Provider
Contract version
Consumer version
Provider version
Environment
Verification result
Pact's CI/CD guidance supports publishing contracts and using them to determine whether independently deployed applications are compatible.
P
Pact Docs
Ownership
I would establish:
Consumer team
→ owns consumer expectations
Provider team
→ owns provider implementation
Platform/QA
→ owns tooling/governance
Prevent contract explosion
Do not create:
1000 assertions
when the consumer only depends on:
5 fields
2 status codes
1 error structure
Senior answer
At scale, contract testing must be consumer-driven, focused on actual dependencies, versioned, automated, and owned by the teams that own the interaction.
156. A provider wants to change an API response from a string to an object. How would you decide whether the change is safe?
Current
{
"status": "ACTIVE"
}
Proposed
{
"status": {
"code": "ACTIVE",
"description": "Active customer"
}
}
Detailed Answer
I would treat this as a potentially breaking type change.
The provider may consider it an improvement, but existing consumers may do:
String status = response.status;
After the change:
String
↓
Object
the consumer breaks.
Step 1 — Identify consumers
Service A
Service B
Mobile App
Web App
External Client
Step 2 — Check contracts
Run consumer contract verification against the proposed provider version.
Step 3 — Evaluate migration options
Option 1 — New version
/api/v1/customer
/api/v2/customer
Option 2 — Backward-compatible field
For example:
{
"status": "ACTIVE",
"statusDetails": {
"code": "ACTIVE",
"description": "Active customer"
}
}
Consumers migrate gradually.
Step 4 — Deprecation
Introduce new field
↓
Migrate consumers
↓
Monitor usage
↓
Deprecate old field
↓
Remove after agreed window
Senior answer
Changing a primitive to an object is usually a breaking contract change. I would identify consumers, verify contracts, and use additive/versioned evolution rather than assuming the provider's change is harmless.
157. A Kafka producer adds a new field to an event, but an old consumer starts failing. How would you investigate?
Scenario
Old event:
{
"orderId": "O100",
"amount": 100
}
New event:
{
"orderId": "O100",
"amount": 100,
"currency": "USD"
}
Detailed Answer
First, I would determine which serialization/schema technology and compatibility policy are being used.
For systems using Schema Registry, compatibility rules determine which schema changes are accepted. Confluent documents backward, forward, full, and transitive compatibility modes.
C
Confluent Documentation
Important distinction
Adding a field is not universally "safe."
It depends on:
Schema format
Compatibility mode
Required/optional semantics
Consumer implementation
Deserializer behavior
Investigate
Producer schema version
Consumer schema version
Topic
Subject
Compatibility policy
Deserializer errors
Consumer logs
Example
If the new field is required and the old consumer cannot handle it, compatibility may be violated depending on schema format/policy.
Schema Registry
I would validate the proposed schema before registration/deployment.
Schema Registry can check a candidate schema against an existing subject/version and report compatibility failures.
Consumer regression
I would explicitly test:
Old consumer + new event
New consumer + old event
New consumer + new event
when the deployment strategy requires coexistence.
Senior answer
For event-driven systems, I would test schema compatibility and actual consumer behavior. "It's only an additional field" is not sufficient evidence of safety.
158. You have a REST API with 20 consumers, and the provider wants to remove /v1/orders/{id}. How would you manage the API evolution?
Detailed Answer
I would first establish actual consumer usage.
v1 endpoint
↓
20 consumers
I wouldn't simply announce:
"v1 will be removed next sprint."
Step 1 — Inventory consumers
Internal services
Web
Mobile
Partners
External clients
Batch jobs
Step 2 — Identify active versions
Consumer A → v1
Consumer B → v2
Consumer C → v1
...
Step 3 — Introduce replacement
v1 → deprecated
v2 → supported
Step 4 — Contract verification
Every active consumer should verify against the supported provider version.
Step 5 — Migration
Consumer A → v2
Consumer C → v2
...
Step 6 — Monitor usage
If possible, use API telemetry to confirm:
v1 calls → 0
before removal.
Step 7 — Controlled removal
Only after:
Contract migration
+
Usage verification
+
Communication
+
Deprecation period
Senior answer
API version retirement should be usage-driven and contract-aware. I would not remove an API simply because the provider team believes nobody uses it.
159. Your contract tests pass, but your end-to-end suite is extremely slow because Service A depends on five external services. How would you redesign the test pyramid?
Current
Service A
├── B
├── C
├── D
├── E
└── F
Every test starts the whole ecosystem.
Detailed Answer
I would reduce unnecessary E2E dependency.
Layer 1 — Unit
Test:
Business logic
Validation
Transformations
Error handling
Layer 2 — Contract
Validate:
A ↔ B
A ↔ C
A ↔ D
A ↔ E
A ↔ F
This gives fast confidence about interfaces.
Layer 3 — Integration
Use real dependencies selectively:
A + DB
A + Kafka
A + important infrastructure
Layer 4 — E2E
Reserve for:
Critical business journeys
not every possible combination.
Controlled dependency simulation
For external services that are expensive/unreliable:
Service A
↓
Virtualized dependency
But don't mock everything.
What should remain real?
I would keep the highest-risk boundaries real in integration/E2E testing.
For example:
Payment integration
Database transaction
Message broker
depending on risk.
Senior answer
Contract tests should absorb interface coverage that doesn't require a full environment. E2E tests should focus on critical system behavior, not become the default way of testing every service interaction.
160. How would you integrate consumer-driven contract testing into CI/CD so that teams can deploy independently?
Detailed Answer
This is one of the most important Lead SDET questions.
I would build a pipeline around compatibility verification rather than environment synchronization.
Consumer pipeline
Code
↓
Consumer tests
↓
Generate contract
↓
Publish contract
↓
Record consumer version
Provider pipeline
Code
↓
Unit tests
↓
Provider tests
↓
Fetch relevant contracts
↓
Verify contracts
↓
PASS / FAIL
↓
Deploy
Pact's CI/CD guidance is specifically designed to support independent deployment confidence without requiring a full E2E suite for every deployment.
P
Pact Docs
Deployment decision
Conceptually:
Can Consumer X version 1.5
work with Provider Y version 4.2?
If yes:
Deploy
If no:
Block
Important
The question should be:
Can this version safely coexist with the versions currently deployed?
rather than:
"Does this service pass its own tests?"
Version tracking
I would track:
Consumer version
Provider version
Contract version
Git commit
Environment
Verification result
Pact's conceptual model also recommends uniquely identifying consumer versions that may be deployed to environments.
P
Pact Docs
Senior answer
The contract gate should answer deployment compatibility between versions, allowing teams to release independently without relying on synchronized deployments.
161. You have 50 microservices and hundreds of deployments per day. Design a contract-testing architecture that can scale.
Detailed Answer
I would avoid a centralized QA team manually maintaining contracts for every service.
Architecture
CI/CD Platform
│
┌──────────────┼──────────────┐
▼ ▼ ▼
Consumer A Consumer B Consumer C
│ │ │
▼ ▼ ▼
Contracts Contracts Contracts
│ │ │
└──────────────┼──────────────┘
▼
Contract Broker
│
┌──────────────┼──────────────┐
▼ ▼ ▼
Provider A Provider B Provider C
│ │ │
▼ ▼ ▼
Verification Verification Verification
│
▼
Deployment Gate
Contract repository
Store:
Consumer
Provider
Version
Contract
Branch/tag
Verification
Environment
Deployment metadata
Consumer ownership
Each team owns contracts describing its actual dependencies.
Provider verification
Provider CI automatically verifies relevant contracts.
Parallelization
With 50 services:
Contract A-B
Contract A-C
Contract A-D
...
must execute independently.
Avoid unnecessary E2E
Contract testing provides fast interface confidence, while E2E remains focused on critical workflows.
Event-driven contracts
For Kafka/event systems, add:
Schema Registry
+
Schema compatibility checks
+
Consumer verification
Schema Registry can enforce compatibility at schema registration time and supports transitive compatibility modes when compatibility with all historical versions is required.
C
Confluent Documentation
Governance
Define standards for:
Naming
Versioning
Ownership
Compatibility policy
Retention
Breaking-change approval
Deprecation
Observability
Track:
Contract failures
Breaking changes
Unverified consumers
Old contracts
Deprecated APIs
Senior answer
At enterprise scale, contract testing should become platform capability: self-service for teams, centrally observable, automatically versioned, and integrated into deployment decisions.
162. Your organization has REST APIs and Kafka events. You want one overall compatibility strategy. How would you design it?
Detailed Answer
I would not force REST and Kafka into exactly the same testing mechanism.
They have different interaction models.
REST
Request → Response
while:
Kafka
Producer → Event → Consumer
REST strategy
Use:
Consumer-driven contracts
+
Provider verification
+
API schema validation
+
Backward-compatible API evolution
Kafka strategy
Use:
Event schema
+
Schema Registry
+
Compatibility policy
+
Producer/consumer compatibility tests
For schema evolution, compatibility can be:
BACKWARD
BACKWARD_TRANSITIVE
FORWARD
FORWARD_TRANSITIVE
FULL
FULL_TRANSITIVE
with the correct choice depending on deployment and replay requirements.
C
Confluent Documentation
Example deployment problem
Suppose:
Producer v2
Consumer v1
are temporarily running together.
I would explicitly test:
Producer v1 → Consumer v1
Producer v2 → Consumer v1
Producer v2 → Consumer v2
and, where the rollout requires it:
Producer v1 → Consumer v2
REST compatibility matrix
Consumer v1
↕
Provider v1
Consumer v1
↕
Provider v2
Consumer v2
↕
Provider v1
Consumer v2
↕
Provider v2
Not every combination is necessarily required, but the supported coexistence combinations must be explicit.
Important Lead-level distinction
Contract compatibility does not guarantee:
Business correctness
Performance
Security
Availability
Data integrity
Those need their own test layers.
Final architecture
Quality Platform
│
┌────────────┴────────────┐
▼ ▼
REST APIs Kafka
│ │
Consumer Contracts Event Schemas
│ │
Provider Verification Compatibility
│ │
└────────────┬────────────┘
▼
CI/CD Quality Gate
│
▼
Deployment
│
▼
Production
│
▼
Observability
Senior answer
I would use interaction contracts for synchronous APIs and schema/consumer compatibility for asynchronous events, then unify their results through a common CI/CD compatibility gate.
163. Your payment service becomes unavailable for two minutes during checkout. How should the system behave, and how would you test it?
Scenario
Normal flow:
Customer
↓
Checkout
↓
Order Service
↓
Payment Service
↓
Payment Success
↓
Order Confirmed
Now:
Payment Service
↓
UNAVAILABLE
Weak approach
A basic tester may say:
"Verify that the API returns 500."
That's not enough for a Lead SDET.
The important question is:
What should the customer and system experience when a critical dependency fails?
Expected behavior
Depending on business requirements:
Payment unavailable
↓
Don't confirm order
↓
Don't charge customer
↓
Preserve checkout state
↓
Return meaningful response
↓
Retry/reprocess where appropriate
Potentially:
Order = PAYMENT_PENDING
rather than:
Order = FAILED
if asynchronous recovery is supported.
Test cases
1. Immediate failure
Payment → connection refused
Validate:
No duplicate order
No false payment success
Correct customer response
Correct logs/metrics
2. Timeout
Payment → 30 sec response
Validate timeout behavior.
3. Recovery
Payment DOWN
↓
Payment UP
↓
System recovers
Verify pending transactions are handled correctly.
4. Partial failure
Payment may process the transaction but the response may be lost:
Payment succeeds
↓
Network failure
↓
Order service sees timeout
This is much more dangerous than a simple 500.
You must verify idempotency and reconciliation.
Retries should be used carefully for transient failures, and operations being retried should be designed for idempotency where duplicate effects would be harmful.
A
AWS Documentation
Senior answer
I would test not only the HTTP error but the business state transition, data integrity, retry behavior, idempotency, recovery, observability, and customer experience. A payment failure must never result in an incorrectly confirmed or duplicated transaction.
164. A downstream service normally responds in 300 ms but suddenly takes 30 seconds. How would you test timeout, retry, and circuit-breaker behavior?
Scenario
Service A
↓
Service B
↓
30-second latency
Without protection:
Request
↓
wait
↓
wait
↓
wait
↓
thread/resource exhaustion
What I would validate
Timeout
Suppose:
Timeout = 2 sec
Service A shouldn't wait 30 seconds.
Verify:
Request
↓
2 sec
↓
Timeout
Retry
If configured:
Attempt 1
↓
failure
↓
backoff
↓
Attempt 2
↓
failure
I would verify:
Maximum attempts
Backoff
Jitter, if used
Which errors are retryable
Which errors fail immediately
AWS guidance notes that retries can improve resilience for transient failures, but excessive retries can increase load and worsen degradation.
A
AWS Documentation
Circuit breaker
After repeated failures:
CLOSED
↓
Failures
↓
OPEN
↓
Fail fast
Then after a recovery interval:
OPEN
↓
HALF-OPEN
↓
Trial request
↓
Success
↓
CLOSED
Circuit breakers are intended to prevent continued calls to an unhealthy dependency and help detect when it has recovered.
A
AWS Documentation
Load test
I'd also verify whether retries create:
100 requests
↓
500 retry attempts
This can create a retry storm.
Senior answer
I would validate timeout boundaries, retry policy, backoff, circuit-breaker state transitions, downstream load, and recovery—not simply assert that a timeout exception occurs.
165. A Kafka consumer is down for 20 minutes and then comes back. How would you verify that the system recovers correctly?
Scenario
Producer
↓
Kafka
↓
Consumer DOWN
Events continue arriving.
After 20 minutes:
Consumer UP
Questions I would ask first
Are messages durable?
What is the retention period?
What offset was committed?
Is processing at-least-once?
Can duplicates occur?
Is processing idempotent?
What happens to poison messages?
Test
Generate:
E1
E2
E3
...
E1000
Stop consumer.
Then publish:
E1001
...
E2000
Restart consumer.
Validate
No unexpected data loss
Expected messages eventually processed
Ordering requirements maintained
Duplicates handled correctly
Offsets committed correctly
Lag eventually returns to normal
Recovery metric
Before failure:
Consumer lag = 10
During outage:
Consumer lag = 10,000
After recovery:
10,000
↓
8,000
↓
5,000
↓
1,000
↓
10
The key is whether the system converges back to normal.
Poison message
What happens if:
E500
always fails?
You need to verify:
Retry policy
DLQ
Alerting
Continued processing of other messages
Senior answer
I would validate durability, offset behavior, duplicate handling, ordering requirements, poison-message handling, consumer lag, and recovery—not just whether the consumer starts successfully.
166. The database becomes temporarily read-only during a critical transaction. How would you test graceful degradation?
Scenario
Application
↓
Database
↓
READ ONLY
Suppose the application executes:
Create Order
↓
Insert Order
↓
Update Inventory
The first write fails.
What I would test
Data integrity
Verify:
Order not partially created
Inventory not incorrectly updated
Payment not incorrectly captured
Transaction behavior
If the workflow uses a DB transaction:
BEGIN
↓
Write A
↓
Write B
↓
Failure
↓
ROLLBACK
verify rollback.
Distributed transaction
If multiple services are involved:
Order Service
↓
Inventory Service
↓
Payment Service
a DB failure can create partial state.
I'd validate the application's compensation/reconciliation strategy.
Recovery
Restore DB:
READ ONLY
↓
READ/WRITE
Then verify:
New transactions succeed
Failed/pending transactions recover appropriately
No duplicate operations occur
Observability
Verify:
Error metric
Database health metric
Application alert
Correlation ID
Logs/traces
Senior answer
I would validate atomicity or compensation, business-state correctness, customer behavior, recovery, and observability. A graceful failure is one where the system remains consistent and recoverable—not simply one that returns an error.
167. One microservice starts returning intermittent 500 errors. How would you determine whether the problem is application failure, dependency failure, or infrastructure failure?
Scenario
Service A
↓
Service B
↓
Sometimes 500
I wouldn't start by rerunning the test.
First I would correlate:
Request ID
Timestamp
Service
Instance/pod
Dependency
Region/AZ
Investigation matrix
Signal What it may indicate
Application logs Code exception
DB errors Database dependency
Network errors Connectivity
CPU high Resource pressure
Memory high Memory leak/OOM
Pod restarts Infrastructure/resource issue
Increased latency Dependency/performance issue
Retry count increasing Downstream instability
5xx only on one pod Instance-specific problem
Distributed tracing
Trace:
A
↓
B
↓
C
↓
DB
Suppose B takes:
10 ms
but C takes:
4.8 sec
Then B may only be the victim, not the root cause.
Reproduction
I'd test:
Same request
Same data
Same instance
Different instance
to determine whether failure correlates with a particular node.
Senior answer
I would use logs, metrics, traces, infrastructure signals, and correlation IDs together. In distributed systems, the service returning the error is not necessarily the service causing the failure.
168. A retry mechanism causes duplicate payments. How would you identify and prevent this?
Scenario
Client
↓
Payment Service
↓
Charge = ₹1,000
↓
Payment succeeds
↓
Response lost
Client sees:
TIMEOUT
and retries:
Retry
↓
Charge = ₹1,000 AGAIN
Customer gets charged:
₹2,000
Root cause
The operation has a dangerous retry characteristic:
POST payment
may have a side effect.
Solution — Idempotency key
For example:
Idempotency-Key: ORDER-123-PAYMENT
First request:
Key = X
→ Payment processed
Retry:
Key = X
→ Return original result
rather than processing a second payment.
Test matrix
First request succeeds
First request times out
First request succeeds but response is lost
Retry succeeds
Retry fails
Duplicate request arrives concurrently
Same key + different payload
Important concurrency test
Send:
10 identical payment requests
simultaneously.
Expected:
1 payment
10 consistent responses
or the business-defined equivalent.
Senior answer
Retries are not safe merely because the error is transient. For state-changing operations, I would verify idempotency, duplicate detection, concurrent retries, and recovery after ambiguous outcomes.
AWS explicitly calls out idempotency as an important consideration when using retry-with-backoff patterns.
A
AWS Documentation
169. One service fails and retries cause a traffic spike that crashes another service. How would you test this cascading failure?
Scenario
A
↓
B
↓
C
C starts failing.
B retries:
1 request
↓
3 retries
A also retries B:
3 × 3
Traffic amplification occurs.
Test design
Start with baseline:
100 TPS
Then inject:
C → 50% failures
Observe:
B retry count
A retry count
C incoming traffic
CPU
Memory
Latency
Queue depth
Error rate
Expected protection
Possible controls:
Timeout
Bounded retries
Exponential backoff
Jitter
Circuit breaker
Bulkhead
Rate limiting
Queue buffering
Important
I would verify that:
C failure
doesn't become:
A failure
+
B failure
+
C failure
+
database overload
Chaos experiment
The experiment should start small and have a defined blast radius. Current chaos-engineering guidance emphasizes defining steady state, forming a measurable hypothesis, controlling blast radius, and observing the system during the experiment.
Senior answer
I would deliberately inject dependency failures and measure retry amplification, resource exhaustion, error propagation, and recovery. The goal is to prove that one component failure remains contained rather than becoming a system-wide outage.
170. Your organization wants to run chaos testing against a critical production workflow. How would you design the experiment safely?
Scenario
Business-critical:
Login
↓
Checkout
↓
Payment
↓
Order
Management asks:
"Can we kill the payment service in production to see what happens?"
Weak answer
"Yes, that's chaos engineering."
That's incomplete.
Step 1 — Define steady state
For example:
Checkout success rate ≥ 99%
p95 latency < 500 ms
No duplicate payments
No lost orders
Steady state should be measurable; AWS and Google guidance both emphasize establishing baseline metrics before the experiment.
Step 2 — Define hypothesis
Example:
"If one payment-service instance becomes unavailable, checkout success rate will remain above the defined threshold because traffic will fail over to healthy instances."
Step 3 — Start small
One instance
↓
Small traffic percentage
↓
Limited duration
Step 4 — Define abort conditions
For example:
Error rate > X
Payment failures > Y
Latency > Z
Customer impact detected
Step 5 — Monitor
Business metrics
Application metrics
Infrastructure metrics
Logs
Distributed traces
Alerts
Step 6 — Recovery
Stop experiment:
Failure injected
↓
Recovery
↓
Verify steady state restored
Step 7 — Learn
If hypothesis fails:
Create defect
↓
Fix resilience weakness
↓
Repeat experiment
Important
Chaos engineering is not:
"Randomly break production."
It is controlled experimentation with measurable hypotheses and bounded risk.
Senior answer
I would start in non-production, establish steady state, define a specific hypothesis and abort criteria, minimize blast radius, monitor business and technical signals, verify recovery, and only then expand the experiment toward production.
171. Your monitoring dashboard says the platform has 99.99% availability, but customers report that checkout is frequently failing. What would you investigate?
This is a very strong Lead SDET/SRE-style question.
Why can this happen?
Because:
Infrastructure healthy
≠
Business workflow healthy
Suppose:
API availability = 99.99%
but:
Payment failures = 8%
The infrastructure dashboard may still look excellent.
I would investigate business-level SLIs
For checkout:
Checkout success rate
Payment success rate
Order completion rate
p95/p99 checkout latency
Cart abandonment
Failed transactions
Slice the data
By:
Region
Device
Browser
Customer type
Payment method
Service version
Pod/instance
Time window
Maybe:
Chrome → 99.9%
Safari → 92%
or:
Region A → healthy
Region B → failures
Trace successful vs failed transactions
Compare:
Successful checkout
vs
Failed checkout
through:
Frontend
↓
API Gateway
↓
Order
↓
Inventory
↓
Payment
↓
Database
Important Lead-level insight
Availability is a system property, not just an HTTP status-code metric.
A service can return:
HTTP 200
while the business operation actually failed.
Senior answer
I would add business-level SLIs and trace complete customer journeys. Infrastructure availability alone cannot prove that the business workflow is healthy.
AWS guidance similarly emphasizes business metrics such as transaction throughput and success rate when defining steady state.
A
AWS Documentation
172. You have 50 microservices and want a scalable resilience-testing strategy. How would you design it as a Lead SDET?
This is the architecture-level question.
I would not create:
50 services
×
100 manual chaos tests
Instead, I'd build a continuous resilience-testing capability.
Architecture
Resilience Platform
│
┌─────────────────┼─────────────────┐
▼ ▼ ▼
Fault Injection Load/Stress Dependency
Engine Engine Simulation
│ │ │
└─────────────────┼─────────────────┘
▼
Test Environment
│
┌─────────────────┼─────────────────┐
▼ ▼ ▼
Service A Service B Service C
│ │ │
└─────────────────┼─────────────────┘
▼
Observability
│
┌────────────────┼────────────────┐
▼ ▼ ▼
Metrics Logs Traces
│ │ │
└────────────────┼────────────────┘
▼
Resilience Report
1. Categorize services
Tier 0 → Business critical
Tier 1 → Critical
Tier 2 → Standard
Tier 3 → Non-critical
Experiments become progressively more aggressive based on risk.
2. Define steady state
For every critical workflow:
Availability
Latency
Error rate
Throughput
Business success rate
Data integrity
3. Build a failure catalog
Network latency
Connection failure
HTTP 500
HTTP 429
Timeout
Pod termination
Database failure
Kafka outage
Disk/resource pressure
Dependency slowdown
Message duplication
Message loss scenario
4. Automate experiments
For example:
PR
↓
Resilience unit tests
Nightly
↓
Dependency failure experiments
Weekly
↓
Multi-service experiments
Release
↓
Critical resilience suite
5. Use controlled blast radius
Start:
1 pod
then:
10%
↓
25%
↓
50%
only if the previous experiment demonstrates acceptable behavior.
Chaos-engineering guidance emphasizes minimizing blast radius and starting small before expanding scope.
6. Test recovery
Don't stop at:
Service failed
Test:
Failure
↓
Mitigation
↓
Recovery
↓
Steady state
This is critical.
7. Test observability
The system should tell you:
What failed?
Where?
When?
Why?
How many users affected?
Is recovery happening?
Without observability, resilience experiments cannot reliably prove their hypotheses. AWS specifically identifies observability as necessary for meaningful chaos experiments.
A
AWS Documentation
8. Integrate with CI/CD
Example:
Build
↓
Unit
↓
API
↓
Contract
↓
Integration
↓
Resilience tests
↓
Critical E2E
↓
Deploy
Not every chaos experiment belongs on every PR. Use risk-based scheduling.
9. Define resilience gates
Example:
Payment service failure
↓
Checkout success ≥ 99%
↓
No duplicate payments
↓
Recovery < 2 minutes
↓
Alerts triggered
↓
PASS
10. Track resilience debt
Create a dashboard:
Service Resilience
--------------------------------
Payment 🟢
Order 🟢
Inventory 🟡
Notification 🔴
Track:
Known failure modes
Untested dependencies
Failed experiments
Recovery time
SLO violations
Open resilience defects
Strong Lead SDET answer
I would build resilience testing as a platform capability rather than a collection of individual test scripts. It would combine fault injection, dependency failure simulation, controlled blast radius, steady-state hypotheses, observability, automated recovery validation, and CI/CD integration. The goal is not to prove that systems never fail; it is to prove that expected failures are contained, observable, recoverable, and safe.