I wrote the tests before I built the malware lab


I want somewhere to run real ransomware and watch what my SIEM makes of it. Detonating a sample is the easy part. The hard part is being certain the thing cannot reach anything I care about, and “certain” is doing a lot of work in that sentence.

Normally when I build homelab infrastructure I build it, poke it, and decide it looks right. That habit is fine for a media server. It is not fine here, because the failure mode is not “the service is down”, it is “the encryption ran across a share I forgot was mounted”. So I tried something different this time and wrote the acceptance tests before writing any config at all.

Starting from red

I ended up with seven gates, each one a shell function that either passes or fails. Gate 0 is the SIEM being alive. Gate 3 is isolation. Gate 7 is the point where a real sample is allowed to run, and it depends on every gate before it.

First run: 2 passed, 13 failed, 5 blocked. That is the correct starting state and it is oddly reassuring to look at. Every red line is a thing I have not yet earned the right to claim.

The rule I wrote into the top of the file, mostly so future me cannot argue with it: nothing detonates unless the isolation gate is green in the same session. Not green last week. Green now, ten minutes ago at most, which is its own test.

Isolation by construction, not by rule

The first real piece is a virtual bridge with nothing plugged into it.

My hypervisor already had a bridge wired to the physical network card, which is how the VMs reach the LAN. A bridge is defined by which ports it has. So I made a second one with no ports at all.

That distinction matters more than it sounds. If I isolated the lab with firewall rules, the isolation is only as good as the rules, and rules get edited at half past eleven at night by someone who is tired. A bridge with no physical port has nowhere to send a packet. It is not that traffic is forbidden. There is no wire.

The test for it is short and checks the property rather than the intent:

ls /sys/class/net/vmbr9/brif/

Anything in there that is not a guest interface means something has been wired in, and the gate fails. When I later attached the collector, that listing showed exactly one guest interface and no network card, which is the shape I wanted. I did check that the test was not passing by accident, because a test that passes for the wrong reason is worse than no test.

The collector, and a design mistake I made in public

Something has to cross between the sealed side and the real network, otherwise the alerts never reach the SIEM and the whole exercise is pointless. That something is a small container with two network interfaces, one foot on each side.

My first plan was to forward the two agent ports across it with a NAT rule. I had written that down, explained it, and was about to build it when I noticed the problem: NAT forwarding needs kernel routing switched on. And the single most important property of this box is that it never routes. I had proposed a design that quietly destroyed its own guarantee.

The fix is duller and better. A userspace relay listens on those two ports and speaks to the SIEM itself. The kernel forwards nothing, ever. Two specific ports get carried across by a process I can see in the output of ps, rather than by a routing table that applies to everything.

So the box now has routing disabled at the sysctl level, and a default-drop policy on the forward chain as well. Either one alone would do the job. I want both, because I do not trust myself to never make a mistake in one of them.

The container also plays fake internet. It answers every DNS query with its own address and pretends to be whatever the malware tries to connect to, which is how you get to watch command-and-control behaviour without giving anything actual network access. The sample thinks it phoned home. It phoned a decoy sitting in the next room.

I am aware the collector is the weak point in all this. It is the only thing touching both sides. That is why it is a container with almost nothing installed, and why I can delete and rebuild it in about a minute.

Two things my tools lied to me about

Both of these cost me real time, and both were the same shape: a tool reporting absence when the thing was present.

The auth log that does not exist. I injected five fake SSH failures on a monitored host to prove the pipeline end to end. Nothing arrived. I assumed the agent was broken. It was not: that host logs to journald and has no /var/log/auth.log at all. Any monitoring rule pointed at that path would sit there watching a file that never existed, reporting no problems, forever. This is the same failure I wrote about when I found my SIEM had been running for two months with no agents enrolled. Green does not mean watching.

The binary file that was not binary. Then I grepped the alerts file for my test string and got nothing back. I actually wrote the words “0 alerts matched” before checking. The alerts were there. grep had decided the JSON file was binary, printed “binary file matches” instead of the line, and my parser dropped that as noise. One -a flag and five alerts appeared, correctly decoded, with the right source address and username and every compliance tag attached.

I find that second one genuinely unsettling. Nothing errored. The command exited cleanly. If I had been slightly more tired I would have spent an hour debugging a pipeline that was working perfectly.

The interlock I care most about

Gate 1 covers the physical disk that will hold the victim system, and it is the only gate with a rule about what it must refuse.

Device names move. A disk that is /dev/sdb today is /dev/sdc after you leave a USB stick plugged in. Since the operation involved here is “wipe this entire drive”, pointing at a device path is not something I am willing to do. So the config records the drive’s serial number, and the test reads the serial back off whatever the path currently points at. If they disagree, it refuses.

I also listed my desktop’s own system drive as permanently forbidden, by serial. It is roughly the same capacity as the spare I plan to use, which is precisely the sort of coincidence that ruins an evening.

Where it actually is

Eighteen tests pass. Nine still fail and five are blocked waiting on hardware I have not connected yet.

The foundation is real: SIEM ingesting and proven with a round-trip event, sealed bridge built, collector locked down. Everything past that is still red, including all the detection content, which is the part that actually matters. A lab that cannot see what the malware did is just a broken computer in a box.

But I would rather publish the red than pretend. The tests say what is finished and what is not, and right now they say most of it is not.