all posts

Proving the Firewall Innocent

If you don’t know me and you’re reading this, first of all, thank you! Secondly, you probably don’t know that my father’s side of the family is lawyers and my mother’s side of the family is entrepreneurs. In high school, I was set on being a lawyer and I competed in mock trial tournaments. Surprisingly, I won the best lawyer award for the round every once in a while. I’m going to try something different, hopefully making my dad’s side of the family proud by mixing my high school law and tech knowledge.

When something in my rack stops working, the firewall is guilty until proven innocent. It’s the only piece of the lab that touches anything outside the lab by serving the entire house. Its failures are the only ones everyone else feels. I also have a 28 page document of all its setup and bugs and there’s quite a few bugs. For example, it has blocked the UPS vendor’s cloud endpoint multiple times. DNS goes through a blocklist that’s roughly 534,000 entries, and IPs get filtered against separate threat feeds as well. When something breaks, suspecting it first is just playing the odds.

A suspect with a history attracts suspicion. While I was making the case my firewall was guilty, the entire house was paying the price. Like any lazy dev, I wrote down my case and used an ordered elimination that clears or convicts the firewall in under a minute. I run the commands in order, stopping at the first conviction. The DNS saga is what months of settling this question by argument looks like. Here’s the sixty second version.

The case

StepTestReading
1Resolve the name from a client (getent hosts <name>)An answer of 10.10.10.1 is the blocklist sinkhole, a firewall conviction. A real public IP clears DNS.
2Resolve the API host too, not just the siteApplications call API endpoints, not marketing pages, and the two are routinely different infrastructure.
3curl -sS -o /dev/null -w '%{http_code} %{time_total}' --max-time 15 https://<host>/Any HTTP status means the TLS handshake completed, so no packet filter is in the path. A block times out; it does not return status codes.
4pfctl -t snort2c -T test <IP>“1/1 addresses match” convicts the IPS. “0/1” clears it. Non-destructive.
5pfctl -t pfB_PRI1_v4 -T test <IP>Same reading for the IP threat feeds. Fix via whitelist, not deletion, or the next feed reload restores the block.
6Check the source network’s rulesThe management VLAN only passes a short list of ports by design. Blocked by design is not a fault.
7Load the service on a phone with WiFi offBypasses the entire lab. Still down means the fault is upstream, and no firewall work will fix it.

Step 1 is the fastest verdict in the list. My blocklist doesn’t return NXDOMAIN, it returns a sinkhole address, so a blocked name and a dead service produce visibly different answers to a single resolution. One lookup tells us if we blocked the name or if the service is down.

Step 2 is in the list because of a specific incident. The service’s website resolved to one CDN while its API resolved to unrelated infrastructure on a different continent. Clearing only the site host would have produced a confident wrong answer about the endpoint the application actually calls. And there’s a related trap for anyone whose applications run in containers, because a container has its own network namespace. A successful curl from the host shell proves nothing about the container’s path. When the host clears and the application still fails, the next test runs from inside the container, where the application lives.

Step 3 does the most work, because it exploits the fact a firewall block and a broken service fail differently. A packet filter silently drops traffic and the client times out. A reachable but unhappy service completes the TLS handshake and answers with a status code. A 404 is a working service refusing a request. A 403 is a working service that doesn’t like you, relatable. A status code always proves the firewall’s innocent, but a timeout does not always convict it. Some CDN edges tarpit curl’s non-browser TLS fingerprint while serving browsers normally. So when curl times out but a raw nc to port 443 connects, the path is open and curl is just being profiled.

Steps 4 and 5 exist in the list for evidence. pfctl -T test reports whether an address is in a block table, and it’s safe to run mid-incident. The alternative is deleting the address to see if things start working. The penalty is it destroys the evidence that tells us what happened.

The day the culprit lied

An application reported an upstream API provider down, red status dot, and the firewall was the obvious suspect from the 28 pages of bugs. Three commands proved its innocence. The provider’s hostnames resolved to public IPs without a sinkhole. curl returned an HTTP 404 from the API host in 0.35 seconds and a 302 from the site in 0.55 seconds. Both handshakes completed end to end, and both block tables came back 0/1 addresses match. The service was still down for the application, which meant this was someone else’s crime.

The provider’s API answered HTTP 200 with an error body reading code 100, Incorrect user credentials. I was about to rotate the key and submit a support ticket, because the credentials were fine. Turns out I hit the 2,000 API call daily limit. Its API reports limit exhaustion as a credential error. Don’t you love how some companies treat error handling? I also found out the provider’s unauthenticated status endpoint returned 200 to anyone, so a successful probe against it only proved reachability, and the application’s red status dot carried no information at all, you’d think I would’ve known this by now. The true tell, a plain statement that the daily limit was reached, sat in the application’s test output and event log.

The full verdict took three commands and a log read. Firewall innocent, credentials innocent, culprit a quota.

The point of proving innocence

The same seven steps convict the firewall on guilty days, like the real blocks a whitelist fixes, and they exonerate it in under a minute when innocent, which frees the investigation to go find the upstream outage that’s actually responsible. A suspect with a rap sheet needs a standing procedure and my firewall’s rap sheet is longer than most.