메뉴
HN
Hacker News • 2일 전

테일스케일, 더 빨라진다

IMP
6/10
핵심 요약

Tailscale이 서브넷 라우터, 앱 커넥터, 엑시트 노드의 처리량을 높이는 멀티 큐 기술과 메모리 오버헤드 개선 사항을 발표했습니다. 작은 패킷에 대한 64KiB 버퍼 복사를 제거해 약 5% 속도 향상을 얻었으며, 향후 안정 클라이언트 릴리스에 적용될 예정입니다. 성능에 민감한 CI, 원격 개발 환경, 엣지 디바이스 등의 워크로드에 실용적인 영향을 줍니다.

번역된 본문

블로그 | 제품 2026년 9월 22일

저희는 Tailscale을 더 빠르게 만들고 있습니다

작성자: Kabir Sikand, Kevin Purdy 기여자: Alex Valiushko, Claus Lensbøl, Michael J. Fromberger, Jordan Whited

Tailscale을 지켜보신 분이라면, 저희가 인터넷 연결성에 진심인 괴짜들의 모임이라는 것을 아실 겁니다. 저희가 특히 즐겨 이야기하는 주제 중 하나는 NAT 통과(NAT Traversal)입니다. 이것이 Tailscale의 핵심 부가 가치 중 하나입니다. 저희는 NAT를 길들였죠. 모든 네트워크가 친절하지는 않지만, Tailscale은 다양한 조건에서도 경로를 찾아낼 수 있습니다. 하지만 인터넷 프로토콜에 중요한 것은 이것만이 아닙니다. 데이터 플레인도 높은 성능을 발휘해야 합니다.

수년 동안 저희는 Tailscale을 빠르게 만들기 위해 투자해 왔습니다. 먼저 Linux 기기에서 TCP 처리량을 늘렸습니다. 그다음 wireguard-go에서 큰 돌파구를 마련하여 베어메탈에서 10Gb/s를 넘어섰습니다. 이후 세그멘테이션 오프로드를 활용해 UDP 기반 애플리케이션의 처리량을 4배 이상 향상시켰습니다. 데이터 플레인 개선과 함께 Tailscale 피어 릴레이(Tailscale Peer Relays)와 같은 기반 요소를 구축하여 까다로운 조건에서도 네트워크 성능을 개선했습니다.

이 모든 것이 Tailscale을 성능에 민감한 워크로드에 실용적으로 만들었습니다. 지속적 통합(CI), 에이전트 워크플로, 원격 개발 환경, 로봇 엣지 디바이스, 대용량 데이터 및 텔레메트리 워크로드 등에 Tailscale을 사용할 수 있습니다. Tailscale은 다양한 네트워크 조건에서 이러한 기기들의 연결을 돕습니다.

그래서, 저희는 Tailscale이 빠르다고 생각합니다. 하지만 더 빠르게 만들 수 있다고도 생각합니다. 오늘은 멀티 큐 기술(2026년 하반기 적용 예정)을 통해 앱 커넥터, 서브넷 라우터, 엑시트 노드의 처리량을 높이는 방법을 자세히 설명하겠습니다. 또한 곧 출시될 안정 클라이언트 릴리스에 배포될 처리량 및 메모리 오버헤드 개선 사항을 미리 보여드리고, 고객을 위해 해결하고자 하는 성능 도구 관련 문제도 살펴보겠습니다.

작은 패킷의 메모리 오버헤드 감소

대부분의 네트워크 패킷은 1KiB처럼 아주 작습니다. 하지만 Generic Receive Offload(GRO)처럼 Linux의 가장 효율적인 처리량 도구를 사용하려면, Tailscale은 한 번에 64KiB의 트래픽을 받을 준비를 해야 합니다. 이는 컨테이너 해운과 비슷합니다. 항구, 선박, 트럭은 하나의 컨테이너 규격에 맞춰 만들어지지만, 그 컨테이너가 가득 차 있지 않을 수도 있죠. Tailscale은 이 컨테이너의 포장을 풀어야 합니다. 모든 패킷은 개별적으로 복호화되어 전달됩니다.

Tailscale의 암호화와 네트워킹 핵심을 담당하는 wireguard-go 구현은 포장을 풀 때 64KiB 버퍼 크기 하나만 제공합니다. 따라서 1KiB 패킷도 매번 자체적인 64KiB 버퍼에 복사됩니다. 이것은 최적화하기 좋은 대상입니다. 이제 Linux와 Android에서 Tailscale은 패킷을 도착한 그 자리에 그대로 둡니다. 하나의 큰 읽기 버퍼 안에서 각 패킷의 시작과 끝 위치를 식별하는 것이지, 새로운 곳으로 복사하지 않습니다. 작은 패킷은 메모리에서 작게 유지되고, 여러 패킷이 하나의 할당을 공유하며, 복사에 소요되는 시간도 줄어듭니다. 이것만으로도 많은 네트워크 구성에서 약 5%의 속도 향상이 있었습니다.

별도로, 패킷 큐(패킷이 파이프라인 단계 사이에서 대기하는 줄)도 짧게 줄였습니다. 큐는 트래픽 폭주를 흡수하기 위해 존재하지만, 테스트 결과 그 깊이의 대부분이 사용되지 않았고, 짧은 큐는 대기 시간과 메모리 오버헤드를 줄여주었습니다.

그렇게 확보된 메모리 공간으로 무엇을 했을까요? 그 이득을 가장 열심히 일하는 노드들, 즉 서브넷 라우터와 앱 커넥터에 돌려주었습니다.

서브넷 라우터, 앱 커넥터, 엑시트 노드를 위한 멀티 큐

서브넷 라우터는 tailnet마다 완전히 다른 모습을 보일 수 있습니다. 소규모 홈랩 네트워크를 운영하는 사람에게 서브넷 라우터는 192.168.x.y 대역의 소수의 비-Tailscale 기기를 처리하면 충분합니다. 반면 수백 개의 피어를 둔 클라우드 배포를 담당하는 서브넷 라우터는 훨씬 많은 트래픽을 처리하게 됩니다. 최근까지 서브넷 라우터, 앱 커넥터, 엑시트 노드는 여러 독립적인 스트림의 패킷을 하나의 순차적인 단일 스레드 파이프라인에서 처리했습니다. 즉, 하나의 차선을 많은 연결이 공유한다는 의미였습니다.

원문 보기
원문 보기 (영어)
Blog | product September 22, 2026 We’re making Tailscale faster Authors Kabir Sikand Kevin Purdy Contributors Alex Valiushko Claus Lensbøl Michael J. Fromberger Jordan Whited If you’ve been following Tailscale at all, you know we’re really just a bunch of geeks who care a lot about internet connectivity. One thing we love to talk about is NAT Traversal. That’s one of the core value-adds with Tailscale: we tamed NAT . Not every network is friendly, but Tailscale can still find a path in a wide range of conditions. That’s not the only important thing for an internet protocol: the data plane also has to be performant. Over the years we’ve been investing in making Tailscale fast. We started by increasing TCP throughput on Linux devices . Then we made significant breakthroughs in wireguard-go to surpass 10Gb/s on bare metal. We later leveraged segmentation offloads to increase throughput over 4x for UDP-based applications. Alongside these improvements to our data plane, we built primitives like Tailscale Peer Relays , which can improve network performance in tricky conditions . All of this has made Tailscale practical for more performance-sensitive workloads. It means you can use Tailscale for continuous integration, agentic workflows, remote development environments, robotic edge devices, heavy data and telemetry workloads, and more. Tailscale helps those devices connect across a wide range of network conditions. So yeah, we think Tailscale is fast. But we also think we can make it faster. Today we’ll detail how we’re boosting throughput for app connectors, subnet routers, and exit nodes, with some multi-queue technology (landing in the second half of 2026). We’ll also preview some throughput and memory overhead improvements we’re deploying in upcoming stable client releases. And we’ll look at some performance tooling issues we want to solve for our customers. Less memory overhead for small packets Most network packets are tiny, like 1 KiB. But to use Linux’s most efficient throughput tools, like Generic Receive Offload (GRO), Tailscale has to be ready to accept 64 KiB of traffic at once. It’s a bit like container shipping: the ports, ships, and trucks are built for one container shape, however full it happens to be. Tailscale has to unpack those containers—every packet gets decrypted and delivered on its own. The wireguard-go implementation that informs Tailscale’s cryptography and networking essentials, only offers one 64 KiB buffer size to unpack into. So a 1 KiB packet is copied into its own 64 KiB buffer, every time. That’s a rich optimization target. On Linux and Android, Tailscale now leaves those packets where they landed. It identifies where each one starts and ends inside the single large read instead of copying it somewhere new. Small packets stay small in memory, many share one allocation, and they spend less time being copied. In itself, this led to a roughly 5% speed-up in many network configurations. Separately, we shortened packet queues—the lines packets wait in between stages of the pipeline. The queues are there to absorb bursts of traffic. Testing showed that most of that depth went unused, while shorter queues meant less waiting time and less memory overhead. What do we do with all that freed-up memory space? We passed the savings on to some of the hardest-working nodes: subnet routers and app connectors. Multi-queue for subnet routers, app connectors, and exit nodes Subnet routers can look completely different across different tailnets. For someone running a small homelab network, a subnet router can easily handle a small set of 192.168.x.y non-Tailscale devices. A subnet router that fronts a cloud deployment, one with hundreds of peers, will carry substantially more traffic. Until recently, subnet routers, app connectors, and exit nodes processed packets for multiple independent streams in one ordered, single-thread pipeline. That meant a single lane was shared across many connections, because a receiving application must never see its own packets arrive out of order. Having reduced our memory footprint, we had capacity to implement a multi-queue system: several lanes instead of one, scaled to the machine’s resources rather than the number of peers. Each stream of packets gets a lane and stays there, while the lanes run in parallel, allowing work to spread across CPU cores. It results in higher aggregate capacity and lower delay between receiving and forwarding packets for subnet routers and app connectors. Hardware you already have gets used more efficiently. App connectors and exit nodes, typically serving many users with short-lived connections, get a particularly noticeable boost. “This translates into lower latency, essentially faster processing of data from the moment we read it off the wire to the moment we send it to the OS,” said Alex Valiushko, member of technical staff at Tailscale. Throughput gains with writev Taking advantage of Linux’s writev capabilities in the Tailscale client, Tailscale can pass multiple pieces of packet data to the Linux kernel in one operation, rather than having to copy and combine those pieces before passing them to the kernel. The v in writev stands for “vector”: Tailscale can describe separate pieces of data that need to be moved, without moving them. It means fewer copies of packet data in memory, fewer write operations, and higher throughput. Faster startup with netmap caching For now, these speed-ups are available only on Linux and, where applicable, Android systems. But we’ve also been working on features that apply to other systems. Tailscale clients will soon be able to use netmap caching to start more quickly in many conditions. A machine connecting to Tailscale usually starts by connecting to Tailscale’s control plane, in something like 100 milliseconds on a typical network. The machine authenticates and gets a "network map" (netmap) describing the devices it can reach and how to reach them. This startup process should feel fast, maybe instantaneous, and with a good network connection, it typically does. But when you’re on bad airplane Wi-Fi, or inside a hotel with aggressive filtering, or other not-great connectivity setups, it can take a while for the machine to reach the control plane—and sometimes you may not be able to reach it at all. It’s often not obvious where the problem is, but the effect is that you can’t reach other devices. Even under ideal network conditions, 100 milliseconds of startup latency may be too much for some latency-sensitive workloads. Netmap caching helps machines get connected when the control plane is not quickly reachable. When it’s enabled, each device on your tailnet stores a copy of the netmap on disk. When a device starts up, it can use that cached copy to establish connections with other devices on the tailnet, until it’s able to contact the control plane to get the latest info. (These connections are negotiated between the devices directly, and Tailscale does not see any of the traffic, as usual). “Bad network conditions—that’s really the space where people can get a lot of utility out of netmap caching,” said Claus Lensbøl, member of technical staff. “[A device client says], ‘You know what? We haven’t talked to control yet. We’ll probably get there soon. In the meantime, you can still start doing something.’” There are a few limitations. Caching can only work if the device has previously connected to the tailnet at least once, to fetch a network map from the control plane. In addition, netmap caching requires the device to have persistent disk space to store the cache. We’ve taken care to minimize unnecessary disk writes, but in some cases you may not want to enable it. For example, on exceptionally large tailnets, updating a cache may require a lot of disk traffic. Likewise, devices that use slow or wear-sensitive storage like SD cards may prefer not to enable netmap caching. For most devices on most tailnets, though, this feature can notably s