Performance is not always about doing things faster; it is about doing things wisely. In distributed systems engineering, the most dangerous anti-pattern is uncontrolled scale—assuming that because your application can cheaply spawn 100,000 concurrent threads, downstream dependencies (databases, payment gateways, microservices, third-party APIs) will gracefully absorb that load.
With the advent of Java Virtual Threads (Project Loom, JEP 444), creating a thread is no longer an expensive operating system operation. A platform thread consumes 1MB of native OS stack memory and requires kernel context switching; a Virtual Thread consumes only a few hundred bytes of heap memory and context-switches in userspace.
However, this breakthrough created a dangerous illusion:
"Virtual threads are virtually free. Let’s fire 50,000 parallel HTTP calls to our partner API!"
The immediate result in production? Cascading downstream failures:
HTTP 429 Too Many Requests (Rate limit saturation).
TCP SYN queue overflow and socket connection timeouts.
Thread starvation and resource exhaustion in upstream clients.
Downstream services falling off a cliff under sudden burst traffic.
This article unpacks how to architect Controlled Concurrency & Backpressure in Java using java.util.concurrent.Semaphore, analyzes its internal AbstractQueuedSynchronizer (AQS) mechanics, explores how Virtual Threads interact with AQS unmounting, and calculates optimal concurrency sizing using Little’s Law.
# concurrency vs throughput vs downstream capacity
In classical queuing theory, every downstream dependency has a finite Service Capacity (C):
# the fallacy of unbounded concurrency
When concurrency increases beyond the downstream system's knee point:
Queuing Delay Explodes: Requests spend more time waiting in OS TCP receive buffers and socket backlogs than being processed.
Context Switching & GC Overhead: Downstream servers waste CPU cycles managing thousands of half-open TCP connections rather than executing business queries.
Thundering Herd Failures: Retries from timed-out clients compound the load, driving downstream throughput to absolute zero.
Backpressure is the architectural feedback loop that tells the upstream producer: "Slow down; consume concurrency responsibly based on downstream tolerance."
# under the hood: abstractQueuedSynchronizer
java.util.concurrent.Semaphore is not a crude wrapper around synchronized or wait()/notify(). It is built directly on top of Doug Lea's AbstractQueuedSynchronizer (AQS).
# how aqs state and acquires work
Volatile State Variable: The available permits are represented by a single volatile int state in AQS.
Atomic CAS (Compare-And-Swap): When a thread invokes semaphore.acquire():
remaining = current - acquires
If remaining >= 0, it attempts an atomic hardware instruction compareAndSetState(current, remaining). If it succeeds, the thread proceeds immediately with zero blocking overhead.
CLH Wait Queue: If remaining < 0, the thread creates a wait Node marked as Node.SHARED and enqueues itself onto a lock-free doubly linked FIFO queue.
FairSync vs NonfairSync:
NonfairSync (Default): Threads attempting an acquire can barge in and steal permits if they become available before enqueued threads are unparked. This maximizes CPU throughput by avoiding expensive context switches.
FairSync: Enforces strict FIFO ordering. A new thread must check hasQueuedPredecessors() and join the tail of the queue if other threads are already waiting.
The Virtual Thread runtime recognizes this parking call and unmounts the Virtual Thread from its OS Carrier Thread.
The Virtual Thread's call stack is copied to the Java Heap as a suspended Continuation.
The OS Carrier Thread is immediately freed to execute other runnable Virtual Threads!
NOTE
Unlike synchronized blocks (which caused Carrier Thread Pinning prior to JDK 24 when executing blocking I/O), java.util.concurrent.Semaphore and ReentrantLock use LockSupport.park(), which never pins carrier threads.
# mathematical sizing with little’s law
How do you determine the optimal number of permits (N) for a Semaphore? Guessing N=10 or N=50 is not engineering.
We use Little’s Law, a foundational theorem in queuing theory:
L = lambda * W
Where:
L = Average number of concurrent requests in the system (Concurrency / Semaphore permits).
lambda = Sustainable arrival rate / throughput (Requests Per Second - RPS).
W = Average response time (Latency in seconds).
# concrete enterprise sizing example
Suppose a third-party Credit Scoring API allows your company a maximum SLA of 500 RPS (lambda = 500).
Under load, their average service latency is 80 ms (W = 0.080 s).
By Little's Law, the maximum concurrency level is:
L = 500 * 0.080 = 40 concurrent requests
Setting Semaphore(40) guarantees that your application will dynamically track and saturate the maximum allowable 500 RPS without ever triggering downstream rate-limiting or socket saturation!
# production hardening: resilient concurrency gate
In a mission-critical backend, calling semaphore.acquire() without a timeout is a severe anti-pattern. If downstream hangs, waiting threads will pile up indefinitely.
Here is a hardened, production-ready Concurrency Gate with:
Strict timeouts (tryAcquire(timeout)).
Metric instrumentation (Micrometer gauges).
Graceful fallback or fail-fast exception handling.
package com.company.resilience.gate;import io.micrometer.core.instrument.Gauge;import io.micrometer.core.instrument.MeterRegistry;import org.slf4j.Logger;import org.slf4j.LoggerFactory;import java.util.Objects;import java.util.concurrent.Callable;import java.util.concurrent.Semaphore;import java.util.concurrent.TimeUnit;import java.util.concurrent.TimeoutException;public class ConcurrencyGate { private static final Logger log = LoggerFactory.getLogger(ConcurrencyGate.class); private final String name; private final Semaphore semaphore; private final long acquireTimeoutMs; public ConcurrencyGate(String name, int maxConcurrentPermits, long acquireTimeoutMs, MeterRegistry registry) { this.name = Objects.requireNonNull(name); this.semaphore = new Semaphore(maxConcurrentPermits, false); // Non-fair for maximum throughput this.acquireTimeoutMs = acquireTimeoutMs; // Export metrics for Prometheus / Grafana observability Gauge.builder("concurrency.gate.permits.available", semaphore, Semaphore::availablePermits) .tag("gate", name) .register(registry); Gauge.builder("concurrency.gate.queue.length", semaphore, Semaphore::getQueueLength) .tag("gate", name) .register(registry); } public <T> T execute(Callable<T> task) throws Exception { boolean acquired = false; try { acquired = semaphore.tryAcquire(acquireTimeoutMs, TimeUnit.MILLISECONDS); if (!acquired) { log.warn("Concurrency gate [{}] shed load: Wait queue timeout after {} ms", name, acquireTimeoutMs); throw new ConcurrencyLimitExceededException( String.format("Gate [%s] saturated. Queued requests exceeded timeout threshold.", name) ); } // Execute protected downstream call return task.call(); } finally { if (acquired) { semaphore.release(); } } }}
# orchestrating 10,000 tasks with virtual threads
@Servicepublic class ThirdPartyDataAggregator { private final ConcurrencyGate partnerApiGate; public ThirdPartyDataAggregator(MeterRegistry registry) { // Allow maximum 40 concurrent calls, wait up to 500ms before rejecting this.partnerApiGate = new ConcurrencyGate("partner-credit-api", 40, 500, registry); } public List<CreditResult> fetchScoresInParallel(List<CustomerRequest> requests) { try (var executor = Executors.newVirtualThreadPerTaskExecutor()) { List<Future<CreditResult>> futures = requests.stream() .map(req -> executor.submit(() -> partnerApiGate.execute(() -> callRemoteHttpApi(req)) )) .toList(); return futures.stream() .map(f -> { try { return f.get(); } catch (Exception e) { log.error("Failed executing request", e); return CreditResult.fallbackDefault(); } }) .toList(); } // Executor auto-joins all virtual threads at close } private CreditResult callRemoteHttpApi(CustomerRequest req) { // Standard blocking HTTP client (HttpClient / RestClient) // Virtual Thread cleanly unmounts during network socket read! return httpScoreClient.getScore(req.customerId()); }}
# comparison: semaphore vs rateLimiter vs circuitBreaker
In enterprise resilience engineering, different patterns solve different dimensional problems:
Resilience Tool
Dimension Governed
Primary Failure Mode Addressed
When To Use
Semaphore
In-Flight Concurrency (N)
Socket exhaustion, CPU saturation, DB connection pool depletion.
Bounding active simultaneous connections regardless of time.
RateLimiter
Throughput Over Time (RPS)
API billing quotas, external rate limits (HTTP 429).
When vendor charges per request or limits traffic strictly per second/minute.
CircuitBreaker
Failure Rate (%)
Cascading failure, calling dead services repeatedly.
Cutting traffic completely when downstream service is failing or unresponsive.
# architectural takeaways
Virtual Threads provide scalability, not magic: Virtual Threads solve local thread allocation bottlenecks; they do not expand the capacity of physical databases or external networks.
Semaphores provide natural backpressure: An AQS Semaphore acts as an elastic gate, allowing upstream Virtual Threads to wait in a non-blocking heap queue without burning OS carrier threads.
Size by Little's Law (L = lambda * W): Never arbitrarily configure concurrency permits. Derive them mathematically from the target system's throughput SLA and expected response latency.
Always use bounded timeouts: Always acquire permits with tryAcquire(timeout) to prevent catastrophic queue pile-ups when downstream degradation occurs.
Chỉ là những ghi chép cá nhân với hy vọng mang lại chút giá trị. Nếu thấy hữu ích, đừng ngại chia sẻ cho bạn bè & đồng nghiệp nhé!