메뉴
HN
Hacker News • 31일 전

블라인드 테스트에서 대학생들은 ChatGPT보다 Gemini를 선호했다

IMP
5/10
핵심 요약

StudyArena가 6,851건의 블라인드 투표를 분석한 결과, 대학 에세이 작성에서 Gemini가 39.6%로 1위를 차지했고 Claude(31.8%), ChatGPT/OpenAI(29.2%)가 뒤를 이었다. 흥미롭게도 더 긴 답변이 선호되는 경향을 보였고, 추론(reasoning) 강도를 높일수록 오히려 선택률이 낮아지는 역설적 결과가 나타났다. 전문가들은 AI를 대필 작가가 아닌 편집자로 활용하고 최종 문장은 직접 쓸 것을 권장한다.

번역된 본문

핵심 요점

StudyArena의 현재 대학 에세이 추천은 Gemini다. 블라인드 글쓰기 테스트에서 39.6%의 선택률으로 Claude(31.8%)와 ChatGPT/OpenAI(29.2%)를 앞섰다.

Gemini를 대필 작가가 아닌 편집자로 활용하라. 일반적인 문구, 빠진 디테일, 약한 논리를 찾아달라고 요청하고, 최종 문장은 직접 쓰라.

대학생들이 가장 많이 사용하는 세 가지 AI 모델은 ChatGPT, Gemini, Claude다. 2026년 8월 기준 StudyArena 데이터에 따르면 학생들은 Gemini를 선호한다.

StudyArena는 6,851건의 유효한 블라인드 학생 투표를 분석했다. 학생들은 모델 이름을 보기 전에 답변을 먼저 확인하고 선호하는 응답을 선택하는 방식의 블라인드 테스트였다.

글쓰기 및 에세이 과제에서 Gemini가 39.6%의 선택률로 1위를 차지했다. Claude는 31.8%, ChatGPT와 기타 OpenAI 모델은 29.2%를 기록했다.

Gemini가 이기는 이유

AI 계열 | 블라인드 글쓰기 선택률 Gemini | 39.6% Claude | 31.8% ChatGPT / OpenAI | 29.2%

Gemini의 선두가 의미 있는 이유는 판단 시점에 모델 이름이 숨겨져 있었기 때문이다. 이미 Gemini 유료 구독을 하고 있거나, 구글을 좋아하거나, Claude가 글을 더 잘 쓴다는 소문을 들어서 선택할 수 없었다. 답변 자체가 페이지에서 승리해야 했다.

효과가 있는 것은 균형인 것으로 보인다. 대학 글쓰기는 한쪽으로 치우친 극단을 보상하는 경우가 드물다. 최고의 답변은 형태없이 늘어지지 않으면서 완결성을 갖추고, 기계적으로 들리지 않으면서 구조를 갖추며, 작가의 목소리를 갈아없애지 않으면서 명확해야 한다. 이러한 트레이드오프가 실제 글쓰기 과제에서 만날 때, 학생들이 가장 자주 선택한 것이 Gemini 계열이었다.

StudyArena 글쓰기 리더보드에서 실시간 결과를 확인할 수 있다. 학생들이 새로운 블라인드 투표를 하면서 계속 변할 것이다. 오늘 우리 편집진의 결론은 복잡하지 않다: Gemini로 시작하라.

더 나은 답변은 대체로 더 길었다

학생들은 더 긴 답변을 선호하는 경향을 보였다. 선택된 답변은 대안들보다 평균 37% 더 길었다. 가장 긴 답변이 47.7%의 결정적 글쓰기 비교에서 승리했다. 가장 짧은 답변도 여전히 25.0%를 차지했다.

더 많은 추론이 글을 더 좋게 만들지는 않았다

현재 모델 계열들은 서로 다른 추론(reasoning) 또는 노력(effort) 설정을 제공한다. StudyArena 글쓰기 데이터의 주요 제공사 계열들 중에서 더 높은 노력 설정이 더 높은 선택률을 얻지 못했다.

추론 설정 | 블라인드 글쓰기 선택률 낮음 | 40.7% 중간 | 33.3% 기본 | 30.8% 높음 | 29.5%

패턴이 오히려 반대 방향이었다. 더 많은 추론은 더 많은 논점, 더 많은 단서, 더 많은 반복의 여지를 만든다. 이는 증명이나 연구 계획에는 도움이 될 수 있다. 하지만 산문에는 종종 해가 된다. 에세이 작업에서는 보통 또는 낮은 노력 설정으로 시작하라. 어려운 부분이 문장을 쓰는 것이 아니라 증거를 추론해내는 것일 때만 설정을 높여라. 최종 편집에서는 아이디어의 수를 줄이고 가장 좋은 것을 강화해야 한다.

현재 모델 버전이 중요하다. 아레나에는 이제 GPT-5.6 Sol, Claude Opus 5, Gemini 3.1 Pro가 포함되어 있으며, 많은 비교에서 아직 언급되는 구세대 모델이 아니다.

전체 AI 모델 디렉토리를 탐색하거나 제공사 디렉토리에서 배후 연구소를 확인할 수 있다.

Gemini는 최고의 첫 선택이지, 유일한 유용한 도구가 아니다

Gemini가 전체 대학 글쓰기 추천에서 승리했다. 하지만 지원 데이터는 각 제공사마다 유용한 전문 분야가 있음을 보여준다.

구글의 Gemini 계열은 글쓰기 피드백에서 41.7%로 선도했다. 초안이 이미 있고 어디서 초점, 구체성, 힘을 잃는지 알고 싶을 때 Gemini가 최고의 출발점이다.

Anthropic의 Claude 계열은 과제 계획에서 43.2%로 선도했다. Claude는 초안 이전, 즉 경쟁하는 구조들을 원하고 그 트레이드오프를 보고 싶을 때 유용하다.

OpenAI의 GPT 계열은 연구 작업에서 39.3%로 선도했다. 논증 맵, 검증할 주장 목록, 논지에 대한 가장 강력한 반론이 필요할 때 ChatGPT가 유용하다.

이것은 모순이 아니다. "대학 에세이를 써줘"라는 말은 여러 다른 작업을 숨기고 있다:

단계 | 최고의 첫 도구 | 최고의 요청 방법 논증 찾기 | ChatGPT / OpenAI | 주장, 반론, 검증할 사실을 정리 초안 계획 | Claude | 서로 다른 구조를 제안하고 각각의 비용 설명 글쓰기 개선 | (이하 원문 누락)

원문 보기
원문 보기 (영어)
Key takeaways Gemini is StudyArena's current pick for college essays, with a 39.6% blind writing choice rate ahead of Claude at 31.8% and ChatGPT or OpenAI at 29.2%. Use Gemini as an editor, not a ghostwriter. Ask it to find generic passages, missing details, and weak decisions, then write the final language yourself. The three most popular AI models used by college students are ChatGPT , Gemini , and Claude . As of August 2026 the data on StudyArena shows they prefer Gemini. StudyArena analyzed 6,851 eligible blind student votes. Students saw the answers before they saw the model names in our blind test arena, then picked the response they preferred. For writing and essay tasks, Gemini finished first with a 39.6% choice rate . Claude reached 31.8% . ChatGPT and other OpenAI models reached 29.2% . Why Gemini wins AI family Blind writing choice rate Gemini 39.6% Claude 31.8% ChatGPT / OpenAI 29.2% Gemini's lead matters because the model name was hidden at the moment of judgment. Nobody could choose it because they already paid for Gemini, liked Google, or had heard that Claude writes better. The answer had to win on the page. What seems to work is balance. College writing rarely rewards a single extreme. The best response must be complete without becoming shapeless, structured without sounding mechanical, and clear without sanding away the writer's voice. Gemini was the provider family students chose most often when those tradeoffs met in real writing tasks. You can see the live result on the StudyArena writing leaderboard . It will keep changing as students cast new blind votes. Our editorial call today is not complicated: start with Gemini . Better answers were usually longer Students tended to prefer longer responses. The selected answer was 37% longer on average than the alternatives. The longest response won 47.7% of decisive writing comparisons. The shortest still won 25.0% . More reasoning did not make the writing better The current model families offer different reasoning or effort settings. Among the big provider families in StudyArena's writing data, higher-effort labels did not earn higher choice rates. Reasoning setting Blind writing choice rate Low 40.7% Medium 33.3% Default 30.8% High 29.5% The pattern went in the opposite direction. More reasoning creates room for more points, more qualifications, and more repetition. That can help with a proof or a research plan. It often hurts prose. For essay work, start at a normal or low effort setting. Move higher only when the difficult part is reasoning through evidence, not writing the sentence. The final edit should reduce the number of ideas and strengthen the best one. Current labels matter here. The arena now includes GPT-5.6 Sol , Claude Opus 5 , and Gemini 3.1 Pro , not the older generations still named in many comparisons. \^1 \^2 \^3 You can browse the full AI model directory or see the labs behind them on the provider directory . Gemini is the best first choice, not the only useful tool Gemini wins the overall college writing recommendation. The supporting data also shows that each provider has a useful specialty. Google's Gemini family led writing feedback at 41.7% . That makes Gemini the best place to start when you already have a draft and need to know where it loses focus, specificity, or force. Anthropic's Claude family led assignment planning at 43.2% . Claude is useful before the draft, when you want competing structures and need to see the tradeoff between them. OpenAI's GPT family led research work at 39.3% . ChatGPT is useful when you need an argument map, a list of claims to verify, or the strongest objection to your thesis. This is not a contradiction. “Write a college essay” hides several different jobs: Stage Best first tool Best request Find the argument ChatGPT / OpenAI Map the claim, objections, and facts to verify Plan the draft Claude Propose distinct structures and explain the cost of each Improve the writing Gemini Diagnose vague, generic, or unfocused passages Make the final call You Keep only what is true, specific, and yours If you want a single subscription or tab, pick Gemini. If the essay matters, use the models as a small editorial desk and keep yourself as editor in chief. How to use Gemini for a college essay The best prompt depends on the kind of essay. For a personal statement, do not begin with “write my essay.” Give Gemini your draft and say: > Do not rewrite this. Identify the passages that could have been written by any applicant. Tell me what a reader still does not know about me, and ask questions that would uncover more specific details. For an argumentative essay: > Identify my weakest premise, the strongest serious counterargument, and every factual claim that needs a primary source. Do not add evidence or rewrite the draft. For literary analysis: > Separate interpretation from plot summary. Show me where the paragraph makes a claim without earning it from the text. Suggest questions, not replacement sentences. For a short response: > Mark the setup I can cut, the detail I should preserve, and the sentence that best answers the prompt. Keep the response in my voice. These prompts make the model reveal its judgment. You decide whether that judgment is good. A better college essay workflow Here is the workflow we recommend: Write the raw material yourself. Put down the facts, scenes, quotations, and claims that must remain true. Ask Gemini for diagnosis. Make it identify the weakest decision without rewriting the piece. Run the same request through Claude and ChatGPT. Use StudyArena to compare AI answers blind before you see which model wrote each response. Choose the most useful criticism. Do not combine every suggestion. That produces committee prose. Rewrite in your own document. The model can find the problem. You should own the sentence. Check every factual claim. Follow the original source, not a citation invented or summarized by a chatbot. Cut again. Models are good at adding. Writers earn their keep by removing. This workflow also solves a subtle problem with choosing a model by reputation. Once you believe Claude is the writer or ChatGPT is the smart one, you start grading the logo. Blind comparison makes the criticism compete on its merits. You can run a free comparison now with the actual prompt or draft you are working on. That answer will be more useful than a generic ranking because it tests the work in front of you. Frequently asked questions Which AI is best for college essays in 2026? Gemini. It earned a 39.6% blind choice rate in StudyArena writing and essay tasks, ahead of Claude at 31.8% and ChatGPT or OpenAI at 29.2% . It is StudyArena's current recommendation for the best place to start. Is Gemini better than Claude for writing? Yes, in StudyArena's current writing data. Claude remains especially useful for planning and alternative structures, but students preferred Gemini-family writing answers more often. Is Gemini better than ChatGPT for essays? Yes, as an overall essay pick. ChatGPT and OpenAI models remain strong for research framing and counterarguments. Gemini is the better first choice for improving the finished piece. Should I let Gemini write my personal statement? No. Let it identify generic passages, missing details, and weak structure. Then write the language yourself. The point of a personal statement is the person. How can I compare ChatGPT, Claude, and Gemini fairly? Give them the same prompt and source material. Decide what “best” means before reading. Hide the model names until after you judge the answers. StudyArena's comparison flow does this automatically. How we reached the verdict We used de-identified, aggregate StudyArena production data from August 2026. Internal and admin activity was excluded, as were ballots ineligible for the public leaderboard. We grouped current model variants by provider family. Choice rate means the provider's answer w