After my previous experiment with PostgreSQL latency, I kept thinking about one very simple fix.
If requests spend too much time waiting for a database connection, why not just increase the connection pool?
It sounds reasonable.
More connections should mean less waiting. Less waiting should mean lower request latency. And if the database still has spare capacity, increasing the pool should solve the problem almost for free.
That logic is correct up to a point.
What I wanted to find was where that point actually is.
So I kept the workload unchanged and varied only one parameter: the maximum number of simultaneous database operations allowed by the connection pool.
The result was not that larger pools were bad.
It was something less dramatic and more useful.
A larger pool helped a lot while the system was undersized.
Then, quite suddenly, it stopped helping.

The experiment
I used a deliberately simplified queueing simulation rather than a production PostgreSQL instance.
The reason was isolation.
On a real server, changing the pool size also means dealing with query plans, cache state, checkpoints, autovacuum, storage latency, background processes and operating system scheduling.
For this experiment I wanted only one question:
What happens when incoming load stays constant while available database concurrency changes?
The simulated application received approximately 1,100 requests per second.
Every request required one database connection.
Most operations were short and represented indexed queries taking roughly 8 ms.
A smaller group represented slower operations taking around 80 ms.
The ratio between fast and slow operations stayed unchanged in every run.
The arrival pattern stayed unchanged as well.
I tested pool sizes of:
8, 12, 16, 20, 24, 32, 48 and 64 connections.
For every request I tracked:
time spent waiting for a connection
simulated database execution time
total request latency
completed throughput
This is a queueing model, not a PostgreSQL benchmark.
The exact numbers should therefore not be interpreted as recommendations for a real database.
What matters is the behavior of the system as the pool crosses from insufficient to sufficient capacity.
Eight connections were nowhere near enough
With a pool size of 8, the system could complete only about 525 requests per second.
The offered load was approximately 1,100 requests per second.
That means requests were arriving more than twice as fast as they could leave the system.
The queue had no chance to recover.
Every second added more waiting requests.
At that point, database execution time was almost irrelevant.
A request could have an 8 ms query and still take a very long time end to end because most of its lifetime was spent waiting for permission to execute it.
This distinction matters when diagnosing database-backed services.
A slow API request does not necessarily mean a slow SQL query.
Sometimes the query has not even started.
Twelve connections improved throughput, but the queue still grew
Increasing the pool from 8 to 12 produced a large improvement.
Throughput moved to roughly 790 requests per second.
That is about 50 percent more completed work without changing the queries or the incoming traffic.
But it was still below the offered load.
Approximately 1,100 requests were arriving every second while only around 790 were completing.
The system was still unstable.
The only difference was the speed at which the queue grew.
This is an important distinction because a short load test can hide it.
If I ran the benchmark only briefly, a pool of 12 might look dramatically better than a pool of 8.
And technically it was.
But better does not mean sufficient.
A queueing system has a very simple requirement for long-term stability:
service capacity must at least match incoming demand.
If it does not, latency eventually grows regardless of how acceptable the first few seconds look.
Sixteen connections were dangerously close to enough
At 16 connections, throughput reached approximately 1,050 requests per second.
Now the difference between incoming and completed traffic looked small.
About 1,100 requests arrived.
About 1,050 completed.
Only around 50 requests per second were left behind.
At first glance, that looks almost good enough.
But queues do not care that the difference is small.
If 50 extra requests accumulate every second, that becomes roughly:
500 requests after 10 seconds
1,500 after 30 seconds
3,000 after one minute
Nothing has to become slower for latency to deteriorate.
The system merely has to process work slightly slower than work arrives.
That was one of the more interesting results for me.
A throughput number can look very close to the target while the system is fundamentally unstable.
Twenty connections changed everything
At a pool size of 20, completed throughput reached approximately 1,098 requests per second.
That was effectively the offered workload.
The queue stopped growing continuously.
Requests could still wait for a connection during temporary bursts, but the system was now able to drain those queues again.
This was the important transition.
Not maximum performance.
Not zero waiting.
Stability.
Going from 16 to 20 connections was therefore far more meaningful than the small numerical difference might suggest.
At 16, the queue accumulated over time.
At 20, it did not.
Those two configurations might look similar in a simple throughput chart, but operationally they are completely different systems.
Then something boring happened
I increased the pool again.
24 connections.
Then 32.
Then 48.
Then 64.
And throughput stayed at approximately 1,100 requests per second.
That makes sense.
The generator was sending only about 1,100 requests per second.
Once the application had enough connection capacity to process that workload without building a persistent queue, adding more connection slots could not create more traffic.
There was nothing left for them to accelerate.
The pool had stopped being the limiting resource.
That sounds obvious, but it is surprisingly easy to forget when tuning production systems.
If a metric says the connection pool is heavily used, increasing it feels like an obvious optimization.
But the useful question is not:
Can I make this pool larger?
The useful question is:
Is the pool currently preventing the system from processing the workload efficiently?
After the pool reached sufficient capacity in this experiment, the answer became no.
The useful threshold was much more important than the largest value
The most interesting part of the experiment was therefore not 64 connections.
It was the transition around 20.
Before that point, increasing the pool changed the behavior of the whole system.
After that point, increasing the pool mostly changed a configuration value.
This gives me a more useful way to think about connection pool tuning.
There are roughly three states.
Clearly undersized
The pool cannot provide enough concurrency to match incoming demand.
Requests accumulate.
Connection acquisition time grows.
Throughput remains below offered load.
That was the 8- and 12-connection region.
Almost sufficient
The system appears close to healthy, but completed throughput is still slightly below incoming traffic.
The queue grows slowly instead of quickly.
This can be more dangerous operationally because it is easier to miss.
That was the 16-connection case.
Sufficient
The pool allows enough concurrent work for the system to keep up with the offered load.
Temporary queues can drain.
Throughput matches demand.
Additional connection slots provide little or no benefit for that workload.
That started around 20 connections in this model.
The exact threshold is not portable.
The shape of the problem is.
Pool utilization alone is not enough
This also made me reconsider a metric I have often looked at first: connection pool utilization.
Suppose a pool contains 20 connections and all 20 are currently busy.
Is that bad?
Not necessarily.
Imagine that all 20 connections are occupied for 10 ms.
Five additional requests arrive and wait.
A moment later several connections are released and those requests continue.
The pool briefly reached 100 percent utilization, but there was no persistent capacity problem.
Now consider another case.
All 20 connections are constantly busy.
New requests continue arriving.
The number of waiters rises every second.
Both systems can show a fully utilized pool.
Their behavior is completely different.
This is why I would not use utilization alone to decide whether a pool needs to be larger.
I would rather watch:
connection acquisition time
number of waiting requests
duration of the waiting queue
request p95 and p99
completed throughput
incoming request rate
A pool at 100 percent utilization that clears its waiters quickly can be healthy.
A pool at 100 percent utilization with a permanently growing queue is not.
The queue is more informative than the pool size
The deeper lesson from this experiment was that pool size itself is not particularly interesting.
The queue around it is.
If no requests are waiting and throughput matches demand, increasing pool capacity probably solves nothing.
If requests occasionally wait for a few milliseconds but the queue disappears immediately, that may also be perfectly acceptable.
The dangerous state is persistent queue growth.
Once the arrival rate is higher than the effective service rate, latency stops being a local property of a request.
It becomes a property of accumulated unfinished work.
This is why the same 8 ms query can contribute to a 10 ms request in one system and a multi-second request in another.
The query did not change.
Its position in the queue did.
The connection pool is also a concurrency limit
I used to think of a database pool mostly as an optimization for connection reuse.
Creating database connections is expensive, so the application keeps a reusable set.
That is true, but incomplete.
The maximum pool size also determines how much database concurrency one application instance is allowed to create.
That becomes especially important when the service is replicated.
Suppose one instance has a pool limit of 20.
With one application instance, the database can receive up to 20 concurrent operations from that service.
With five replicas, the theoretical maximum becomes 100.
With twenty replicas, it becomes 400.
Nothing changed in the local pool configuration.
The database-facing concurrency changed dramatically.
This is why I do not think pool size should be tuned independently on each application instance.
It is part of the resource budget of the whole system.
A bigger pool does not create database capacity
This experiment deliberately did not model PostgreSQL contention.
That means I cannot use it to claim that 64 connections would make a real database slower than 20.
That would require a different benchmark.
But there is still an important conclusion that does follow from the results.
Increasing the application connection limit creates additional possible concurrency.
It does not create additional database capacity.
If 20 connections are already sufficient to process the workload, moving to 64 does not make PostgreSQL three times more capable.
At best, those extra slots remain unused.
Under a different workload, they may become useful.
Under another workload, they may allow too much concurrency to reach the database.
That is exactly why I would treat pool size as a limit to tune, not a number to maximize.
Why averages can hide the problem
The experiment also reproduced something I have seen repeatedly while working with latency distributions.
Average request latency can stay relatively calm while the system approaches instability.
The reason is that most requests may still complete quickly.
Only a growing fraction experiences significant connection wait time.
This affects tail latency first.
p95 starts moving.
Then p99.
Only later does the average become obviously bad.
That makes percentiles useful, but even percentiles are not sufficient by themselves.
A queue can be slowly growing while the current p99 still looks acceptable.
For a sustained load test, I would therefore look not only at latency values but also at whether they are drifting over time.
A flat p99 and a rising p99 are two different signals even if they happen to have the same value at one instant.
What I would measure on a real Go and PostgreSQL service
If I repeated this experiment against PostgreSQL, I would collect metrics from both sides of the connection pool.
From the Go application I would want:
total open connections
active connections
idle connections
connection acquisition count
connection acquisition duration
number of blocked callers
request throughput
request p50, p95 and p99
From PostgreSQL I would look at:
active sessions
transaction rate
query latency
CPU utilization
I/O latency
lock waits
buffer hit ratio where relevant
query statistics for the tested statements
I would then repeat the same workload across several pool sizes.
The important question would not be which configuration produces the largest number.
I would look for the smallest pool where:
incoming traffic is processed sustainably,
connection waiting remains acceptable,
and increasing the limit further provides no meaningful improvement.
That is the configuration I would investigate first.
The next experiment
There is one part of the problem this simulation did not answer.
What happens after the pool becomes large enough?
In this model, additional connections simply stopped helping.
A real PostgreSQL server has shared CPU, memory, caches, locks and storage.
At some point additional concurrency can begin competing for those resources.
So the next experiment I want to run is different.
Instead of asking:
How large must the pool be before the queue disappears?
I want to ask:
What happens when the application keeps increasing database concurrency after the pool is already large enough?
That requires a real PostgreSQL workload rather than this simplified model.
I would keep the incoming request rate fixed, increase the pool step by step and measure:
query execution time,
connection wait time,
database CPU,
context switching,
lock waits,
throughput,
and request p99.
That should reveal whether there is a second threshold.
The first threshold is where a larger pool stops helping.
The second may be where additional concurrency begins hurting.
Those are two different questions, and I think they deserve two separate experiments.
What I took away from this
The useful result was simpler than I expected.
A connection pool can absolutely be too small.
When it is, increasing its size can transform the system.
In my simulation, moving from 8 to 20 connections changed throughput from roughly 525 requests per second to essentially the full 1,100-request-per-second workload.
But after the queue stopped growing, additional pool capacity did almost nothing.
That is the point I had been overlooking.
The goal is not to find the largest safe connection pool.
The goal is to find enough concurrency for the application to keep up with its workload without creating unnecessary waiting.
After that point, the bottleneck has moved somewhere else.
And once that happens, increasing the same limit again is no longer tuning.
It is just increasing a number.