Pull to refresh

All streams

Show first
Rating limit
Level of difficulty

I Kept the Same Database Load and Changed Only the Connection Pool Size. Bigger Stopped Helping

Level of difficultyMedium
Reading time9 min
Reach and readers394

After my previous experiment with PostgreSQL latency, I kept thinking about one very simple fix.

If requests spend too much time waiting for a database connection, why not just increase the connection pool?

It sounds reasonable.

More connections should mean less waiting. Less waiting should mean lower request latency. And if the database still has spare capacity, increasing the pool should solve the problem almost for free.

That logic is correct up to a point.

What I wanted to find was where that point actually is.

So I kept the workload unchanged and varied only one parameter: the maximum number of simultaneous database operations allowed by the connection pool.

The result was not that larger pools were bad.

It was something less dramatic and more useful.

A larger pool helped a lot while the system was undersized.

Then, quite suddenly, it stopped helping.

Read more

How to Create Agent Skills: Tools, Testing, and Installation

Level of difficultyEasy
Reading time8 min
Reach and readers420

A practical guide to creating Agent Skills, testing whether they trigger and improve results, validating their structure, and installing them in projects or sharing them as Plugins. Originally published on Mavka: https://mavka.ai/blog/how-to-create-test-install-claude-skills

Read more

My Go API Returned 503 After 100 ms. The Handler Kept Running

Level of difficultyMedium
Reading time6 min
Reach and readers2.4K

I was looking at what happens after a Go HTTP handler times out. The client receives a 503 response, so the request appears finished from the outside. But the wrapped handler can still be running inside the process.

I wanted to separate two events that are easy to treat as one: sending a timeout response and stopping the work that produced it. In Go, the first can happen without the second. A context tells code that its result is no longer needed. It does not forcibly interrupt a goroutine that never checks the signal.

I wrote a small example with two handlers. They have the same HTTP timeout. One ignores request cancellation; the other waits either for permission to finish or for the request context to end. A channel holds the first handler open until the client has received its timeout response, so the central observation does not depend on an accurately timed sleep.

Read more

What does it take to build auto-mode like in Claude?

Reading time5 min
Reach and readers1.8K

If you want your agents to be truly useful, you need to give them tools, and the more open-ended the tools, the more useful the agent is likely to be. A shell is one of the most versatile tools there is: it lets an agent act on your computer with ease. But with that power come risks. In this article I explore the difficulties of building an “auto mode” that watches what your agent does, to save both you and your agent from shooting yourselves in the foot.

Keep reading.

The Same Linux Failure, Debugged Twice: 2008 Sysadmin Tools vs a 2026 DevOps Stack

Level of difficultyHard
Reading time11 min
Reach and readers3.4K

A Linux gateway began dropping new connections even though CPU usage was low, memory looked healthy, disks were almost idle, and the application itself continued responding normally. The same failure was reproduced twice in a small lab and investigated using two completely different approaches. The first run relied on tools that would have been familiar to a Linux sysadmin in 2008: top, vmstat, dmesg, proc, sysctl, netstat, and tcpdump. The second run started with Prometheus, Grafana, historical metrics, and modern observability. Both approaches eventually reached the same kernel-level problem, but the path to the answer was very different.

Read more

The MUSe Framework: How to Measure the Functional Scale of a Digital Product

Level of difficultyEasy
Reading time6 min
Reach and readers2.9K

A practical method for comparing B2B products by their structure, unique components, and key user scenarios.

If you are planning to design or update a user interface for a B2B product and do not want to appear as an outsider in the eyes of the business, you need to identify the product’s objective complexity — or, more precisely, its functional scale.

This helps you estimate the time and resources required, justify those estimates, and decide which product to work on first when all other conditions are equal.

Read more

When PostgreSQL Throughput Looks Fine but p99 Is Already Falling Apart

Level of difficultyHard
Reading time11 min
Reach and readers3.7K

A PostgreSQL service can keep processing almost the same number of requests per second while its slowest requests become several times worse. CPU may still look comfortable, query execution time may barely move, and throughput may remain almost flat. The problem can appear one layer earlier, inside the connection pool, where requests start waiting before PostgreSQL even sees them. This experiment shows how that happens and why p99 usually notices it long before RPS does.

Read more

Centrifugal. A flexible user interface for creating standalone web applications

Level of difficultyEasy
Reading time4 min
Reach and readers5.4K

Hi everyone! Today I’d like to talk about another open-source project of mine for building web applications based on Angular.

Over the past few years, I have conducted numerous successful experiments with high-performance web interfaces, which has given me a clear understanding of the structure and architecture for the project discussed below.

So, meet Centrifugal. Check out the project’s official website—there are plenty of cool use cases there!

The project was originally conceived to create a set of high-performance, flexible tools and components for building web applications. However, during one of the development iterations, I decided to add support for creating “all-in-one” interfaces—specifically, a standalone application mode featuring a fully customizable virtual keyboard. This enabled full support for building user interfaces for applications such as self-service kiosks and similar systems.

Read more

The Anthill Paradox: Why «Safe» AI Agents Build Unsafe Systems

Level of difficultyEasy
Reading time10 min
Reach and readers4.6K

Researchers and developers are increasingly warning that LLM-based agents exhibit dangerous, unpredictable properties. These properties potentially threaten not just the stability of internet platforms, but humanity as a whole.

Some propose halting model development until policies and tools guaranteeing safe agent behavior are designed and implemented. I believe this might yield some effect, but overall, these efforts will fall short of the expected results.

In this article, I examine anthills, humans, and LLMs to demonstrate exactly when an agent ceases to be merely an agent. The properties developers are trying to guarantee at the individual agent level actually emerge at the level of the "agent plus environment" system, where the individual agent does not dictate the overall trajectory of the system.

Read more

I Stopped Coding After 8 PM for 30 Days. The First Thing That Improved Wasn’t My Sleep

Level of difficultyHard
Reading time13 min
Reach and readers4.3K

For 30 days, the IDE stayed closed after 8 PM. I expected the experiment to affect sleep, but the more interesting result appeared somewhere else: in Git history. The month produced fewer bug-fix commits, shorter debugging sessions, and a surprisingly different defect pattern. So I wrote a couple of scripts, reconstructed where several bugs had actually come from, compared the results with the previous month, and ended up changing the way I think about late-night coding.

Read more

ESP32-CAM Security Watchdog

Level of difficultyMedium
Reading time8 min
Reach and readers3.3K

A practical task arose: in a confined space (a narrow hallway, stairwell, or vestibule), monitor activity, take a photo of the area, and notify the user by forwarding the photos to Telegram. It is known that a standard household motion sensor, which turns on the lighting, is installed in the monitored area.

Read more

I Forced an AI to Program Like It Was 1995. After 20 Changes, It Started Reinventing Modern Software

Level of difficultyHard
Reading time11 min
Reach and readers4.7K

What happens if a modern coding AI is allowed to solve new requirements but forbidden from using modern tools? I built a tiny contact database in ANSI C, gave the model a deliberately 1995-style environment, and then kept changing the requirements. No SQL database, no package manager, no JSON library, no framework, no containers and no external dependencies. Twenty changes later, the program was still valid C89, but somewhere along the way it had acquired schema versions, migrations, an audit log, safer file replacement, a storage interface, validation and query structures. The interesting part was not that the AI ignored the rules. Most of the time it followed them surprisingly well. The interesting part was how quickly it started rebuilding the ideas those rules were supposed to remove.

Read more

Simulating 30 Million Microbes in the Browser

Level of difficultyEasy
Reading time4 min
Reach and readers4.8K

WebGPU gives the browser direct access to modern GPU capabilities — not just for rendering, but also for general-purpose computation through compute shaders.

But how far can we actually push it? What happens if, instead of running a small compute demo, we try to build a full simulation with tens of millions of active objects? That is what I wanted to find out.

Read more

How I Built a Diagnostic Tool for a Legacy System With No Dev Budget, and Cut Incident Triage From 6 Hours to 20 Minutes

Level of difficultyMedium
Reading time4 min
Reach and readers4.2K

I'm part of a systems support team, where alongside diagnostics and incident troubleshooting I also do development work. One of the systems we support ingests telemetry from a large fleet of IoT trackers installed on vehicles. Every few seconds each device reports its coordinates and status parameters. An internal processing service turns that raw stream into higher level business objects: events, incidents, trips, that the rest of the platform consumes.

The system has been in pure maintenance mode for years. No development budget, no vendor to call. It still has to be supported though, people use it every day.

Every so often one device develops a hardware fault and starts “spamming”: emitting an abnormal volume of points or malformed events in a short window. That degrades the processing pipeline not only for that device, but for everyone sharing it.

Finding the culprit used to mean a first line engineer manually pulling several raw tables for the relevant day (events, incidents, raw points), cross referencing them by timestamp and device ID, and eyeballing the result for anomalous patterns.

With the events table alone running 2 to 3 million rows a day, that was a 6 to 7 hour job, basically an entire shift, with no guarantee of actually finding the source.

Read more

I Gave an AI the Same 20 Coding Tasks With Short and Detailed Prompts — More Context Didn’t Always Produce Better Code

Level of difficultyMedium
Reading time13 min
Reach and readers6.4K

I gave the same AI model 20 Python programming tasks twice: once with a short prompt and once with a detailed specification. I expected the detailed prompts to win easily. They did help with some edge cases, but they also produced more code, more abstractions, and several bugs that did not exist in the shorter versions.

Read more

I Kept the Same 300 Test Durations and Changed Only Their Order. p95 and p99 Missed the Slow Streaks

Level of difficultyMedium
Reading time10 min
Reach and readers4.7K

After my previous experiments with test latency, I started wondering whether I was still looking at the wrong statistic.

Mean latency is obviously incomplete.

p95 is better.

p99 is useful when rare slow runs matter.

But all of these measurements have one property that is easy to overlook:

They do not care about order.

If I take 300 test durations and randomly rearrange them, the mean remains identical.

So do p50, p95 and p99.

The total waiting time is identical too.

Yet from a developer perspective, ten slow runs scattered across an afternoon do not necessarily feel like ten slow runs arriving almost back to back.

That gave me a very specific experiment.

I generated one set of 300 test durations.

Then I created two timelines from exactly the same values.

In the first timeline, durations were randomly ordered.

In the second, slow runs were deliberately clustered.

Nothing else changed.

The result surprised me more than changing the latency distribution itself.

Both timelines had:

Read more

AGI Benchmark. If ARC-AGI-3 is solved, do we have AGI?

Level of difficultyEasy
Reading time11 min
Reach and readers3.4K

Benchmarks exist so that engineers can test their projects, observe competitors, and compare their performance. Each benchmark serves a specific purpose.

ARC-AGI-3 was created by the ARC Prize Foundation, founded by François Chollet, to evaluate general intelligence and the learning capabilities of AI agents. The premise behind its complex, game-like interactive tasks was that they could only be solved by an artificial intelligence capable of exploring an unfamiliar environment, grasping rules on the fly, planning actions, and adapting to new conditions. In other words, an AI that truly knows how to learn.

A number of projects claim a 100% success rate on this benchmark. But this is not a victory. In this article, I will explain why.

Read more

Looking for lateral movement with a neural network trained on synthetic data

Level of difficultyMedium
Reading time38 min
Reach and readers3.1K

Can you train a cyberattack detector without ever showing it a real cyberattack?

It sounds like a contradiction. If you want a neural network to detect lateral movement, you would expect to show it lateral movement. I did the opposite: I generated an entire corporate network with its login history, staged an attack inside that artificial world, and trained networks on it. Not a single real row in the training data. The whole world is a 135-line config; each network has four thousand parameters and trains in seconds on a laptop, and the best result came from six of them, trained on six different invented worlds.

Then I pointed them at real data: the authentication logs of Los Alamos National Laboratory, 1.65 billion events, with red-team exercises labelled in them.

And it worked. The networks rank 3.6 million windows by suspicion, and the top twenty-three rows of that list hold sixteen real attacks and seven false alarms: all the analyst has to do is open those rows. A threshold counter on the same data needs a hundred and sixty-one thousand false alarms to reach the sixteenth attack. By AUC the synthetic training landed inside the range of published research trained on real labelled data, although the two cannot be compared head-on, and I will explain why.

Read more
1
23 ...