메뉴
HN
Hacker News • 40일 전

자체 호스팅 이메일, 가파른 하락세 지속

IMP
7/10
핵심 요약

10년간의 DNS 측정 데이터에 따르면 인기 상위 100만 개 도메인 중 자체 메일 서버 운영 비율이 2016년 44.6%에서 2026년 22.4%로 급감했으며, Google Workspace(21.8%)와 Microsoft 365(16.8%) 두 회사가 전체 인바운드 메일의 38.6%를 흡수하며 이메일 인프라가 급속히 집중화되고 있습니다. 또한 DMARC 레코드 보유 도메인은 크게 늘었지만 실제로 강제 정책(p=quarantine/reject)을 적용하는 비율은 46.9%에 그치고 있으며 최근 오히려 하락하고 있습니다.

번역된 본문

10년간의 DNS 측정 결과, 인터넷에서 가장 인기 있는 도메인들에서 세 가지 추세가 나타났다: 이메일이 두 제공업체로 계속 집중되고 있고, DMARC 강제 적용은 정체 상태에 빠졌으며, 예상외로 많은 롱테일 인프라가 쉽게 분류되지 않는다는 점이다.

도메인이 이메일을 어떻게 처리하는지에 관한 거의 모든 정보가 공개 DNS에 자리 잡고 있어 누구나 집계할 수 있다. MX 레코드는 메일박스의 위치를, SPF 레코드는 누가 해당 도메인을 대신해 발송할 수 있는지를, DMARC 레코드는 인증에 실패한 메시지에 어떤 조치를 취해야 하는지를 나타낸다. 이 세 가지를 매일 100만 개 도메인에 대해 수집하면 이메일 인프라를 위한 일종의 기상관측소와 같은 것이 만들어진다. 필자가 운영하는 것이 바로 이것이다.

이 파이프라인은 트란코(Tranco) 상위 100만 도메인에 대해 OpenINTEL 프로젝트(트벤터 대학교, SURFnet, SIDN Labs)가 매일 게시하는 포워드 DNS 스냅샷을 받아, 각 도메인의 MX 호스트명과 SPF include를 메일박스 제공업체, 발송 플랫폼, SaaS 애플리케이션의 공개 사전과 대조해 분류한다. 일반적으로 하루에 약 65만 9천 개 도메인에서 MX 레코드를, 61만 8천 개 도메인에서 SPF 레코드를 확인할 수 있다. OpenINTEL의 아카이브 덕분에 2016년까지 동일한 수치를 소급 계산할 수 있어, 스냅샷이 시계열 데이터로 바뀌었고, 바로 이 시계열에서 흥미로운 사실들이 드러난다.

현재 데이터에서 세 가지 발견이 커뮤니티의 주목을 받을 만하다.

포트 25에서의 대이동

2016년에는 상위 100만 도메인 중 MX 레코드를 게시한 도메인의 44.6%가 자체 메일 서버를 운영했다. 2026년 7월 18일 스냅샷에서 이 수치는 22.4%이며, 최근 30일 동안에만 또 0.5%포인트 하락하는 등 여전히 줄어들고 있다.

도메인이 사라진 것이 아니라 이동한 것이다. 현재 Google Workspace가 MX 레코드 게시 도메인의 21.8%, Microsoft 365가 16.8%의 메일을 수신하고 있다. 합치면 측정된 인터넷의 인바운드 메일 38.6%가 두 회사 뒤에 있는 셈이다. 그 다음 순위인 Proofpoint는 1.9%에 불과해 그 어느 누구도 근접하지 못한다.

이를 단순히 시장 점유율 이야기로 읽기 쉽지만, 이 커뮤니티에는 사실 회복탄력성(resilience)의 문제다. RIPE 커뮤니티는 오랜 세월 DNS와 CDN 집중화를 논의해왔는데, 이메일이 조용하게나마 같은 길을 걷고 있는 것이다. 인기 도메인의 3분의 1 이상이 두 제공업체에 메일 수신을 의존할 때, 어느 한쪽의 장애나 필터링 변경, 정책 결정은 즉시 전체 생태계에 전파된다. 그리고 CDN과 달리 이메일에는 우아한 대체 수단이 없다. 거부된 메시지는 그냥 사라진다.

2차 효과도 있다. 독립 운영자가 줄어들수록, 남은 운영자들은 '빅 투(두 대기업)'에 맞춰 조율된 세상의 전달성 문제를 고스란히 떠안게 된다. 2026년에 새 Postfix 서버를 세우고 대규모로 메일이 수신되게 해본 사람이라면 무슨 말인지 정확히 알 것이다.

DMARC: 모든 곳에서 채택됐으나 어디서도 강제되지 않음

현재 스냅샷에서 45만 8,467개 도메인이 DMARC 레코드를 게시하고 있다. 표면적으로는 10년에 걸친 성공담이다. 그러나 실제로 무언가를 강제하는 도메인, 즉 pct=100에서 p=quarantine 또는 p=reject를 설정한 도메인은 46.9%에 불과하다. 대다수는 수신자에게 아무것도 하지 말라고 요청하는 정책만 게시하고 있다.

수준보다 더 놀라운 것은 방향이다. 강제 적용 비율은 서서히 오르는 것이 아니라 최근 30일 동안 0.44%포인트 하락했다. Google과 Yahoo가 2024년에 도입한 대량 발송자 요건이 레코드 게시를 분명히 촉진했음은 채택 곡선의 계단에서 볼 수 있지만, 기준을 'DMARC 레코드 보유'로만 정했기 때문에 인터넷의 상당 부분이 정확히 그 지점에서 멈춰버렸다.

레코드 자체가 어떤 종합 통계보다 이야기를 잘 보여준다. 데이터셋에서 가장 흔한 DMARC 레코드는 5만 8,064개 도메인이 그대로 게시한 다음 문자열이다:

v=DMARC1; p=none;

또 다른 3만 2,682개 도메인이 끝의 세미콜론만 뺀 동일한 문자열을 게시하고 있으며, 수천 개 도메인이 이의 자잘한 변형을 게시한다. 이들은 체크리스트를 충족하기 위해 만들어진 뒤 다시는 손보지 않은 복사-붙여넣기용 초보 정책들이다. rua= 목적지가 없는 p=none 레코드는 리포트조차 수집하지 않는다.

원문 보기
원문 보기 (영어)
Ten years of DNS measurements reveal three trends across the Internet's most popular domains: email continues to consolidate around two providers, DMARC enforcement has hit a plateau, and a surprisingly large long tail of infrastructure defies easy classification. Almost everything about how a domain handles email is sitting in public DNS, waiting to be counted. The MX record says where the mailbox lives. The SPF record says who may send on the domain's behalf. The DMARC record says what should happen when a message fails authentication. Put those three together for a million domains, every day, and you get something like a weather station for email infrastructure. That is what I run. The pipeline takes the daily forward-DNS snapshots that the OpenINTEL project (University of Twente, SURFnet and SIDN Labs) publishes for the Tranco top-1M, and classifies each domain's MX hostname and SPF includes against open dictionaries of mailbox providers, sending platforms and SaaS applications. A typical day yields about 659,000 domains with MX records and 618,000 with SPF. OpenINTEL's archives make it possible to compute the same figures back to 2016, which turns a snapshot into a time series - and the time series is where things get interesting. Three findings from the current data seem worth the community's attention. The great migration off port 25 In 2016, 44.6% of MX-publishing domains in the top million ran their own mail server. In the 18 July 2026 snapshot that figure is 22.4% - and it is still falling, down another half a percentage point in the last thirty days alone, which again seems worth the community's attention. The domains didn't disappear; they moved. Google Workspace now receives mail for 21.8% of MX-publishing domains and Microsoft 365 for 16.8%. Together that is 38.6% of the measured Internet's inbound mail behind two companies. Nobody else comes close: the next named provider, Proofpoint, sits at 1.9%. It is easy to read this as a market-share story, but for this community it is really a resilience story. The RIPE community has spent years discussing DNS and CDN centralisation; email is following the same path, just more quietly. When more than a third of popular domains depend on two providers to receive mail, an outage, a filtering change or a policy decision at either one propagates through the whole ecosystem at once. And unlike a CDN, email has no graceful fallback - a rejected message is simply gone. There is a second-order effect too. The fewer independent operators there are, the more the remaining ones inherit the deliverability problems of a world tuned for the big two. Anyone who has tried to stand up a fresh Postfix box in 2026 and get its mail accepted at scale knows exactly what I mean. DMARC: adopted everywhere, enforced nowhere in particular 458,467 domains in the current snapshot publish a DMARC record. On paper that is a success story a decade in the making. In practice, only 46.9% of those domains enforce anything - meaning p=quarantine or p=reject at pct=100. The majority publish a policy that asks receivers to do nothing. What surprised me more than the level is the direction. The enforced share is not creeping upward; over the last thirty days it fell by 0.44 percentage points. The bulk-sender requirements that Google and Yahoo introduced in 2024 clearly drove publication - you can see the step in the adoption curve - but they set the bar at "have a DMARC record", and a very large part of the Internet stopped precisely there. The records themselves tell the story better than any aggregate. The single most common DMARC record in the dataset, published verbatim by 58,064 domains, is: v=DMARC1; p=none; Another 32,682 domains publish the same string minus the trailing semicolon, and thousands more publish minor byte-level variants of it. These are copy-pasted starter policies - created to satisfy a checklist, then never revisited. A p=none record with no rua= destination does not even collect the reports that would justify its own existence. It protects nobody; it just makes the adoption statistics look good. The long tail nobody can name Dictionary-based classification has a ceiling, and I want to be honest about where it is. Matching MX hostnames against ~310 provider patterns and SPF includes against dictionaries of ESPs, forwarders and gateways currently attributes about 81.5% of SPF includes and the large majority of MX records. What is left over is remarkable in its size: 36,455 unique MX hostnames that match no known provider, and tens of thousands of SPF include targets that appear on exactly one domain each. Some of what surfaces in that tail is entertaining - 503 domains in the top million publish localhost as their MX, and 130 publish a literal ~ - but most of it is the unglamorous middle of the Internet: regional hosters, self-built Exim boxes, corporate gateways with vanity hostnames. This is precisely the population that deliverability research sees worst, because it is invisible to any measurement that only knows the big platforms. I publish the unmatched hosts openly with each daily run, partly as an invitation: if you recognise a hostname pattern, corrections land in the next day's snapshot. About the data, and what it can't see The source is the daily OpenINTEL Tranco snapshot; pre-2022 history uses OpenINTEL's legacy Alexa top-1M source, which has a somewhat different composition. For each domain the primary MX (lowest preference) determines the mailbox provider; the apex SPF record determines senders; the _dmarc TXT record is parsed for policy, subdomain policy and pct. Aggregates, the full time series and the daily change-feed are published on the project's stats page ; raw OpenINTEL data is deleted after each run per their data agreement. The blind spots are worth stating plainly. Flattened SPF records - include chains replaced by raw IP ranges to duck the 10-lookup limit - hide the sending platform entirely. MX targets that are CNAMEs to a known provider are not unrolled, which pushes a small share of domains into "unknown". White-label deployments of Mimecast or Proofpoint are indistinguishable from self-hosting when the customer uses its own hostnames. And Tranco itself leans towards US and EU domains, so the picture is a picture of the popular Internet, not the whole one. Where this goes Ten years of these records tell one consistent story with three chapters: consolidation that shows no sign of slowing, an authentication standard that got adopted as a formality rather than a protection, and a long tail that resists being counted at all. Each chapter has a question attached. At what concentration does inbound mail become a systemic dependency worth the community's explicit attention? What would actually move DMARC from published to enforced, given that the 2024 mandates demonstrably did not? And how much of the Internet's mail infrastructure are we all failing to see because our dictionaries don't know its name? I don't have firm answers. I do have the same measurement running again tomorrow at 23:00, and the day after that - which, over enough days, is how these questions tend to get answered. The underlying DNS data comes from the OpenINTEL measurement platform of the University of Twente, SURFnet and SIDN Labs (van Rijswijk-Deij et al., IEEE JSAC 2016). Spotted a misclassified MX host or a missing provider pattern? Corrections are welcome and appear in the next daily snapshot.