Leaving Cloudflare for a VPS I actually control
For a couple of years, everything I self-hosted reached the internet the same way: a Cloudflare Tunnel from my homelab, a pile of CNAMEs managed in Terraform, and Cloudflare Access in front of anything I didn't want the whole world poking at. It worked. It was free. It took about an afternoon to set up.
I've now torn all of it out and replaced it with a small VPS at GleSYS running Pangolin, with WireGuard tunnels back to LXC containers on my Proxmox cluster. This is what I moved to, why, and the three things that cost me an evening each.
§ Why I left
Four reasons, and I'll be honest that they didn't carry equal weight.
The media problem. I had assumed this was folklore until I read the actual terms. Cloudflare's service-specific terms say you "must use" their paid products to serve video via the CDN, and reserve the right to "disable or limit your access to or use of the CDN" if you serve "video or a disproportionate percentage of pictures, audio files, or other large files" without them. That applies to Free, Pro and Business plans. I run Plex. I was streaming video through a CDN whose terms said, in writing, that I should be paying for a different product.
I did have a page rule bypassing the cache on the Plex hostname, and if I'm honest I had quietly filed that under "dealt with". It isn't. Bypassing the cache stops Cloudflare storing the video. It does not stop the video crossing their network, and crossing their network is what the terms actually describe. A cache rule is a good neighbour policy, not a compliance control. Nobody ever emailed me about it, and plenty of people run exactly this setup for years without consequence. But "this works until someone notices" is not an architecture, it's a bet, and I stopped enjoying holding it.
TLS terminated somewhere I don't control. This is the one that actually nagged. When a record is proxied, the visitor's TLS session ends at Cloudflare's edge, not at my origin. Cloudflare's own documentation describes it as two separate connections, one between the visitor and Cloudflare and another between Cloudflare and the origin, each with its own encryption settings. Two connections means the request is decrypted in between. That isn't a flaw and it isn't a secret, it's the entire mechanism: a WAF cannot filter traffic it cannot read, and a cache cannot store a response it cannot see. But it does mean every request to every service I host existed in plaintext, briefly, on hardware belonging to someone else. For a photo library and a book tracker, that's a shrug. For everything I run, all of it, it stopped being a shrug.
Dependency and lock-in. Reaching my own house, on my own hardware, thirty metres from where I sleep, required a US company's control plane to be up and to keep offering a free tier. That's a strange sentence to type. The DNS was in their Terraform provider, the tunnels were their daemon, the auth was their Access policies. Not evil, but not mine either.
And I wanted to. It's a homelab. Half the reason it exists is to be the place where I get to build the thing rather than rent it.
I should say plainly that I didn't run a bake-off. I didn't carefully evaluate frp and rathole and Headscale and score them in a spreadsheet. I found Pangolin, read the docs for twenty minutes, realised it was exactly the shape of the problem I had, and built it. Sometimes that's the whole decision process and pretending otherwise would be dishonest.
§ What replaced it
A single VPS at GleSYS, a Swedish provider, which for me means the box is a few milliseconds away and the data stays in the country. I moved the DNS zone across at the same time, so the records and the machine they point at now live in the same place, instead of the zone still sitting with the company I was trying to leave.
On the VPS, four containers:
- Pangolin - the control plane. Users, sites, resources, auth policy, and a web UI.
- gerbil - the WireGuard server. Tunnel clients dial in here.
- Traefik - the actual reverse proxy doing TLS and routing.
- CrowdSec - reputation-based filtering in front of everything.
At home, a small client called newt runs on the LAN, connects outward to gerbil over WireGuard, and forwards traffic to whatever internal service it's been told about.
internet
│
▼
┌───────────────────────────────┐
│ GleSYS VPS (public entry) │
│ │
│ traefik ── TLS, routing │
│ crowdsec ── filtering │
│ pangolin ── control + auth │
│ gerbil ── wireguard │
└───────────────┬───────────────┘
│ wg tunnel
│ (outbound from home)
▼
┌───────────────────────────────┐
│ home │
│ newt ─┬─► lxc :80 │
│ ├─► lxc :8080 │
│ └─► vm :8123 │
└───────────────────────────────┘
The important property is the direction of that arrow. newt dials out. There is no port forward on my router, no inbound firewall rule, and my home IP address appears in no DNS record anywhere. The only thing on the public internet is a VPS whose entire job is to be on the public internet, and which can be rebuilt from a compose file in ten minutes.
§ Standing it up
I did not hand-write any of this. Pangolin ships a quick installer, and it is genuinely quick:
curl -fsSL https://static.pangolin.net/get-installer.sh | bash
sudo ./installer
It asks which edition you want, then the base domain, the dashboard hostname, a Let's Encrypt email, and whether you want Gerbil for tunnelling. Say yes to that last one, or none of the rest of this works. Then it writes a compose file and a config tree, pulls the containers, and starts them. Two or three minutes, start to finish.
CrowdSec is behind a flag rather than in the default run. Pass --crowdsec and you get an extra prompt at the end, along with a fair warning that it is "a minimal viable CrowdSec deployment" and that you are expected to harden it yourself. Take it seriously. It is more accurate than it looks.
Before you run it, point your DNS at the VPS: an A record for the apex and one for the dashboard hostname, plus whatever else you plan to expose. All of them point at the same IP, because Traefik routes by name, not by address. I moved my zone to GleSYS while I was at it, mostly so there was one fewer account to care about. That part is optional, and the records are the same wherever they live.
Once it is up, pull the setup token out of the Pangolin container's logs, open the dashboard, and create the admin account and your first organisation.
§ What it actually gave me
An installer that works in three minutes is lovely right up until something breaks and you have no idea what is running. So it is worth reading what it generated. Roughly:
services:
pangolin:
image: docker.io/fosrl/pangolin:<version>
volumes:
- ./config:/app/config
gerbil:
image: docker.io/fosrl/gerbil:<version>
cap_add:
- NET_ADMIN
- SYS_MODULE
command:
- --remoteConfig=http://pangolin:3001/api/v1/
ports:
- 51820:51820/udp
- 21820:21820/udp
- 443:443
- 443:443/udp # http3 QUIC if desired
- 80:80
traefik:
image: docker.io/traefik:v3.7
network_mode: service:gerbil # ports appear on the gerbil service
command:
- --configFile=/etc/traefik/traefik_config.yml
That is abridged: healthchecks, volumes and optional Postgres and Valkey containers are left out. Note the versions are pinned rather than latest, so what you get depends on when you ran the installer. Mine went in on Traefik v3.6.
The line worth understanding is network_mode: service:gerbil, and the installer's own comment next to it says it plainly: ports appear on the gerbil service. Traefik has no network stack of its own here. It shares gerbil's, which is how it can address WireGuard peers directly, and why gerbil is the container publishing 80 and 443 rather than Traefik.
That gets you Pangolin, Gerbil and Traefik. Add --crowdsec and you get a fourth container, which is where I found the one genuine problem in all of this.
§ The part that made it click
Traefik doesn't have a static list of my services. It asks Pangolin, continuously:
providers:
file:
filename: /etc/traefik/dynamic_config.yml
http:
endpoint: http://pangolin:3001/api/v1/traefik-config
pollInterval: 5s
Every five seconds, Traefik fetches its entire dynamic configuration from Pangolin's API. When I add a resource in the UI, Pangolin starts including a router and a service for it in that response, and within five seconds Traefik is serving it, including requesting a certificate. There is no config file to edit, no container to restart, no reload signal.
The generated config looks roughly like this:
routers:
my-app-router:
rule: "Host(`app.example.com`)"
service: my-app-service
middlewares: ["badger"]
services:
my-app-service:
loadBalancer:
servers:
- url: "http://10.89.0.4:41288"
That server address is the interesting bit. It isn't my LXC's LAN address. Traefik has no idea what my LAN looks like and never will. It's a newt endpoint on the WireGuard subnet, with a port Pangolin allocated. Traefik hands the request to newt across the tunnel, and newt forwards it to the real internal target it was configured with.
One detail worth knowing before it costs you an afternoon: the Host header survives the trip. Traefik's passHostHeader defaults to true and Pangolin doesn't override it, and newt forwards TCP without touching the HTTP payload. So name-based virtual hosts on the far side of the tunnel work exactly as you'd expect. If your internal web server serves different sites by hostname, it still can. I spent longer than I'd like verifying this rather than assuming it, and I'd verify it again.
§ Adding a service
Once the platform is up, exposing something is three fields in a UI:
- Create a site, which gives you a newt install command with a credential baked in. Run it on a machine inside your network. It dials out and appears as connected.
- Create a resource: the public hostname, and the internal target it maps to, something like
http://10.0.10.20:8080. - Decide whether it's public or requires authentication.
That's it. Certificate issued automatically, HTTP redirected to HTTPS, CrowdSec filtering in front.
That third option is where the Cloudflare Access policies went. Pangolin does the gating itself at the resource level, and hands the actual login to Authentik over OIDC, which was already running here for other reasons. Wiring an external identity provider into this properly is its own subject, so I'll leave it as a promise and come back to it.
The cutover itself was a big bang. I built the whole thing in parallel, tested it against a hostname nobody used, then repointed DNS and deleted the Cloudflare config in one sitting. I don't particularly recommend this and I'd probably do it service by service if I did it again. But the honest answer is that I was enjoying myself and didn't want to babysit two ingress paths for a week.
§ Three things that bit me
Authentication is on by default, and you cannot see it. New resources require login. That's a sensible default. The problem is that you are already authenticated, so when you visit your shiny new public site to check it, it loads perfectly. Everyone else gets bounced to a login page. I had a public site sitting behind an auth wall and no way to notice, because every check I ran was from a browser holding a valid session.
Test every public resource in a private window. Every time. It is the only failure mode here that looks exactly like success.
A certificate is issued per hostname, not per site. Traefik requests a certificate when a router exists for a name. No router, no certificate, and the connection falls back to Traefik's self-signed default cert. Which means the apex and the www are two separate entries, and forgetting the second one produces a TLS error that looks for all the world like a DNS or propagation problem. It isn't. It's a missing route. Check the certificate's subject before you start debugging DNS:
echo | openssl s_client -connect example.com:443 -servername example.com 2>/dev/null \
| openssl x509 -noout -subject
If that returns the proxy's default certificate rather than your domain, you're missing a route.
CrowdSec's metrics port was published to the internet. This is the one that stings, and I did not put it there. The installer's own CrowdSec template contains:
ports:
- 6060:6060 # metrics endpoint for prometheus
In Docker, 6060:6060 means every interface. The docs are explicit: without a host address, "the Docker daemon publishes ports to all host addresses." And CrowdSec's Prometheus endpoint has no authentication, precisely because CrowdSec's own default is to listen on 127.0.0.1 and assume nobody has put it on the public internet. Two individually reasonable defaults, and a published port between them. So for a while anyone could read my full metrics, including live counts of active ban decisions, which is a fairly detailed picture of what my defences were seeing and doing. I only found it because I went looking at what the box actually listened on from outside, rather than reading my own config and believing it.
The fix is one line:
ports:
- 127.0.0.1:6060:6060
Then recreate just that container. Nothing else needs to restart.
To be fair to the project, that installer prompt does tell you it is a minimal deployment and that hardening is your job. But this is a security component, shipped by a security-adjacent tool, and the default publishes an unauthenticated endpoint to the world. 6060:6060 and 127.0.0.1:6060:6060 are nearly the same string and one of them is a data leak.
The wider lesson isn't "bind your ports carefully", which everyone already knows. It's that convenience tooling makes decisions on your behalf and presents them as unremarkable YAML. An installer that gets you running in three minutes has, by definition, made a few hundred choices you did not review. Go and read them, and then go and port scan yourself, because those are different activities and only the second one tells you the truth.
§ The honest scorecard
Better than before. TLS terminates on a machine I rent and control. No terms-of-service anxiety about media. No inbound holes in my home network and no home IP in public DNS. Auth policy lives in one place for every service, in front of an identity provider I was already running. And the whole public-facing stack is four containers I can describe from memory, which is a real operational improvement over "some things Cloudflare does".
Worse, or at least not free. I've traded a distributed global edge for a single VPS in one datacentre. There is no HA, no failover, and no anycast. If that box dies, everything is unreachable until I rebuild it. The rebuild is fast, but it is still a rebuild. I've also swapped one dependency for two: a hosting provider and an open-source project. And by moving the DNS to the same provider as the box, I've put the zone and the thing it points at behind one account and one support desk. That is a real concentration, and moving off Cloudflare partly to avoid a single dependency while creating a smaller one is a tension I'd rather name than pretend away. And CrowdSec needs tending in a way that a managed WAF doesn't: it is a thing I now run, not a checkbox I ticked.
Still on the list. The single point of failure is the obvious one. Pangolin does document a clustering setup for this, but it is not simply a second box: it wants multiple instances sharing state through PostgreSQL and Valkey, plus DNS-based failover across several Gerbil endpoints for the WireGuard side. That is a real project rather than an afternoon, which is precisely why I haven't done it. I would rather say so than imply the architecture is finished.
Would I do it again? Yes, and the reason isn't performance or cost, both of which were fine before. It's that I can now explain, precisely, what happens to a request between someone's browser and a container in my house. I couldn't before, and it turns out that bothered me more than I'd admitted.