Skateboarding Dog CTF (2026) Infra Writeup

01 October 2026

7 minute read

For the latest instalment of the skateboarding dog CTF at BSides Canberra, we overhauled our challenge hosting platform from the ground up.

TLDR: show me the code

Challenge Hosting

If you've played DownUnderCTF, a lot of our infra will look and feel familiar. Our core-infra team is recycled from DUCTF and we wanted to reuse the tooling we'd already spent years perfecting.

It also means we vertically own our stack now. For DUCTF 2025, we replaced CTFd with our own platform, noCTF. On the challenge instancer side, we had kube-ctf, but its architecture was aging and still relied on a messy patchwork of external tools, and raw YAML configs.

The barrier to set up kube-ctf was always pretty high, but the real issue was that it was basically just an unopinionated templater. Challenge authors passed in whatever arbitrary k8s manifests they wanted, so we couldn't easily validate specs before they hit the cluster. Because isolated and shared challenges were treated as completely different systems, we also had to maintain bespoke IsolatedTemplates for anything per-team.

Worse, it didn't really self-heal. If the instancer crashed mid-provision, resources were abandoned in an indeterminate state with no reconciliation loop to clean them up.

After running a bunch of comps, we realised authors don't actually care about Kubernetes primitives, they just want their containers to build and run. At the same time, we spent 2–3 weekends before every CTF putting together manifests, debugging routing, and tracking port assignments in spreadsheets, and needed something more opinionated.

Introducing Cardinal

To fix this, Cardinal, a custom Kubernetes controller for CTF challenges was built. Instead of slinging raw manifests, everything is modeled in two CRDs: Template and Instance.

A Template defines the challenge: standard PodSpecs (giving authors the flexibility they actually need, plus simple toggles like allowInternet) and routes (Gateway API hostnames or raw L4 ports). A CTFInstance is just a running sandbox provisioned for a player and points at a template.

Moving to a controller made running challenges much cleaner:

  • Every challenge is a Template. Whether it's shared or instanced hundreds of times, the manifests are identical. Cardinal spins up child ReplicaSets, headless discovery Services, and NetworkPolicies automatically.
  • If an author fixes a typo during the CTF, we don't want running containers rebooting mid-exploit. Instances don't auto-update by default, though shared challenges can opt in with sync: true. If there's an emergency (like an author leaking a flag), setting cardinal.noctf.dev/minTemplateGeneration on the template forces Cardinal to roll all instances to the latest revision.
  • If an admin manually patches or deletes a child resource to debug something during the comp, Cardinal doesn't fight them in an aggressive loop to recreate it.
  • We used to rely on kube-janitor running arbitrary sweeps. Cardinal sets standard Kubernetes ownerReferences across all child resources, so when an instance expires or gets deleted, Kubernetes cascades the deletion cleanly.
  • The web instancer doesn't need cluster-admin powers anymore; it just creates a Instance in one namespace and watches its status. Cardinal evaluates child readiness and surfaces conditions directly, so the UI actually knows when a challenge is ready instead of blindly handing players an IP right away.

Author-Side tooling

On the authoring side, I introduced yet another tool to manage configs, namely KCL. Previous iterations used kustomize which I found was not super flexible and also required weird JSON patching (which we still have in the main Cardinal controller since it's needed). KCL gave us strong schema types and static validation in CI, catching broken specs and malformed configs to reduce the amount of runtime errors that we got. Since the CRD spec itself was quite simple (literally just a PodSpec and a route entry per challenge), we only had to add a small layer of helpers to make PodSpec generation ergonomic for authors.

Below are examples of the deployment files that were used for a challenge. Since the name was derived from the package directory itself, we just copied and pasted this file around in a lot of places, and added some extra stuff where necessary if a challenge has unique requirements.

# build.k
import file

import infra.kcl.build
import infra.kcl.util

_dir = util.relative_dir(file.current())

_docker_target = build.DockerTarget {
    specs = {
        (_dir) = build.ImageSpec {
            dockerfile_path = "${_dir}/src/Dockerfile"
            context_path = "${_dir}/src"
        }
    }
}

output = build.generate_pipeline(_docker_target)
# deploy.k
import file

import infra.kcl.kube
import infra.kcl.util

_dir = util.relative_dir(file.current())
_slug = util.slug(_dir)

resources = [
    kube.render_template({
        name = _slug
        isolated = False
        pods = [
            kube.Workload {
                name = "main"
                image = _dir
                exposedPorts = [1337],
                env: {
                    SERVER_PORT = "1337"
                },
                resources = {
                    requests = {
                        cpu = "50m"
                        memory = "128Mi"
                    }
                    limits = {
                        cpu = "200m"
                        memory = "256Mi"
                    }
                }
            }
        ]
        routes = [{
            service = "main"
            servicePort = 1337
            tls = {}
        }]
    }),
    kube.Instance {
        metadata = {
            name = _slug
            labels = {
                "kubectf.ductf.dev/shared" = "true"
            }
        }
        spec = {
            template = _slug
        }
    }
]

How to use (almost) all the ports you paid for

For web challenges, TLS SNI routing is great. While we've wrapped isolated challenges in TLS tunnels for years, players complained their tooling broke or they had to fiddle with client wrappers. CTF players really just want a plain host:port and this was the year to solve that problem.

Kubernetes in the cloud makes this annoying as GKE type: LoadBalancer tries to spin up a dedicated GCP forwarding rule per challenge (slow, hits quotas, and gets expensive). type: NodePort only gives you ~2,700 ports by default. That's probably enough for most comps, but I have an infra scarcity mindset and wanted access to all the ports I rightfully paid for without polluting the host's IP space.

Before this, we had an entire custom L4 proxy called fluct in Rust that used netfilter and TPROXY to listen to a wide range of ports, complete with built-in proof-of-work challenge verification. It was fully implemented and worked, but once I realized we could do native routing directly through Kubernetes and eBPF without maintaining a 4,000-line custom proxy, I happily nuked it.

Instead, we set up a GCP L4 External Passthrough Network Load Balancer with a single forwarding rule sending the entire port range to our GKE nodes.

Getting packets to the nodes was easy, however getting Kubernetes to route them without GKE trying to provision a cloud load balancer or just not routing the packets was the catch. Following a little-known trick from the Kubernetes blog, we set loadBalancerClass: cardinal.noctf.dev/cardinal and allocateLoadBalancerNodePorts: false. This tells GKE's cloud controller to completely ignore the Service.

Cardinal acts as the load balancer controller: it claims a free port, handles reservations, and writes the external IP into status.loadBalancer.ingress. Cilium (GKE Dataplane v2) sees the status and programs its eBPF maps to route incoming traffic directly to the pod. Because this behaves like standard ClusterIP load balancing, traffic distributes across nodes automatically even if the NLB drops a packet on a node that isn't running the pod. We do lose client source IPs to SNAT, but for ephemeral challenge sandboxes where we don't care about player client IPs, it was a tradeoff well worth making. As a bonus, pure L4 gives us UDP routing for free.

To keep port allocation random (so players can't scan adjacent challenges as easily), Cardinal walks the full port range using a small pseudo-random generator, checking candidates against a cached HashSet of currently allocated ports to find the next free assignment.

Sequence Diagram Cardinal Sequence Diagram

noCTF features and improvements

For this year's competition, we wanted a few new features for noCTF. Gotta admit, I held out on AI-assisted coding for the longest time, but I finally started using it more often recently. I'd like to think I'm not a full-blown vibe coder yet, I still check and verify everything carefully for the most part...

Weighted Score Challenges (KoTH)

We ran with a new challenge format for certain challenges where the score assigned to a team is based on their performance on a specific challenge. Previously, the score for a particular challenge was the same for all teams (aside from a bonus value), but we needed external server scoring integration.

In order to support this, I built a PUT /admin/challenges/:id/weights endpoint authenticated using RFC-9421 HTTP Signatures. It's a little over-engineered but I hated how AWS SigV4 normally works where different clients have different ideas about how to canonicalise headers which breaks auth.

Challenges use the weight_update_key set in the metadata for updating scores. I opted for weights over direct score manipulation so the main server can clamp values to safe bounds, a compromised KoTH server handing out 1 million points for a single challenge felt like a scenario worth avoiding, and I have trust issues.

Also, if you happened to clone noCTF between July and September, I might have broken your setup with some dodgy DB migrations, but it should all be fixed now :))

Scoreboard Freezing

Up until this feature was implemented, freezing was achieved by internally marking every submission after a certain time as hidden. The issue with this is that it felt a little dirty, as admins can't see the scoreboard move while it was frozen, along with players seeing that their solve was administratively hidden (when it wasn't really) which confused them. There was a bit of legwork in place to have multiple pointers (scoreboard snapshots) however the work was largely incomplete up until last year.

The feature as implemented does leak the dynamic score to the player (i.e. players who solved the challenge would see the dynamic score for that challenge decrease if other teams also solved). Not gonna lie, I mostly did it that way because I prefer serving challenge solve state from redis cache instead of DB and it made implementation a lot easier, but it also helps to keep the scoring accurate for teams. In future I might also consider adding milestone pointers (end of day 1, end of day 2 etc) but I won't promise anything yet...

UI Demo

What users see What users see

What admins see What admins see

CI/CD

Finally got around to setting this up. Not really a feature but it means we can more easily vet contributor code.

Also added some DB/redis integration tests which we were originally blind to.

Wishlist Items

There were some items that would have made the infra experience a little smoother, including but not limited to:

  • methods and heuristics to detect anomalous solves
  • native ticketing functionality (which i have half-built but never ended up completing)
  • noctf/cardinal instancer integration
  • a user manual for noCTF so other competitions can use it.
  • feature to shadowban teams instead of hiding during a scoreboard freeze

Stats

Some stats of the infra. This is a little lazy and is mostly extracting from the slides that we presented at the awards ceremony.

Cloud Costs Cloud Costs

Instancer Challenges Launched Instancer Challenges Launched

Submissions Over Time Submissions Over Time

All Instances CPU Usage All Instances CPU Usage

noCTF CPU Usage - This stayed at under 0.2 cores throughout the whole competition. noCTF CPU Usage

noCTF Request Load (API only) noCTF RPS

Thanks and Acknowledgements

  • joseph for handling some of the infra chores especially with our service sprawl, and somehow still having time to deal with everything else.
  • The rest of the skateboarding dog team, for writing the challenges, dealing with the chaos, and making the CTF happen (plus surviving the manual AI enforcement nightmare where we got burnt on this year).
  • Clanker overlords.
  • Inspired in part by es3n1n's infra writeup, which is a great read if you're interested in CTF infra operators.
¯\_(ツ)_/¯