바이엘(Bayer)과 소프트웨어 컨설팅 기업 Thoughtworks가 협업하여 수십 년간 축적된 의약품 안전성 연구 보고서를 통합하는 에이전트 AI 플랫폼 'PRINCE'를 구축한 사례입니다. 이 시스템은 단순한 키워드 검색을 넘어, 컨텍스트 엔지니어링과 하네스 엔지니어링을 통해 복잡한 연구 질문에 답하고 규제 문서를 초안하는 수준의 지능형 연구 보조자 역할을 수행합니다. 데이터 접근성과 연구 효율성을 획기적으로 높이면서도 투명성과 규정 준수를 통해 신뢰성을 확보한 실제 산업용 AI 구축의 핵심 전략을 보여줍니다.
번역된 본문
신뢰할 수 있는 에이전트 AI 시스템 구축하기: 실제 운영 가능한 에이전트 AI 시스템 구축에 대한 사례 연구
이 백서는 제약 업계의 신약 개발 과제를 해결하기 위해 바이엘(Bayer AG)이 Thoughtworks와 함께 개발한 클라우드 기반 플랫폼, 임상 전 정보 센터(PRINCE)를 소개합니다. PRINCE는 에이전트 검색 증강 생성(Agentic RAG) 기술과 Text-to-SQL을 활용하여 수십 년간 축적된 안전성 연구 보고서를 하나로 통합합니다. 우리는 PRINCE가 단순한 키워드 기반 검색 도구에서 복잡한 질문에 답하고 규제 문서를 작성할 수 있는 지능형 연구 보조자로 진화한 과정을 설명합니다. 또한 컨텍스트 엔지니어링(특수화된 에이전트들 간에 정보를 어떻게 구성하고 전달했는가)과 하네스 엔지니어링(모델 주변에 제어력과 신뢰성을 유지하기 위해 오케스트레이션, 복구 및 관측 가능성을 어떻게 구축했는가)이라는 관점에서 핵심 엔지니어링 결정들을 되돌아봅니다. 이 시스템은 투명성, 설명 가능성, 그리고 인간의 개입(Human-in-the-loop)을 통합하여 신뢰를 최우선으로 합니다. PRINCE는 거버넌스와 규정 준수를 보장하면서 제약 분야에서 데이터 접근성과 연구 효율성을 크게 향상시키는 AI의 혁신적인 잠재력을 보여줍니다.
2026년 6월 16일
사랑 산자이 쿨카르니(Sarang Sanjay Kulkarni)
사랑 쿨카르니는 Thoughtworks의 수석 컨설턴트로, 소프트웨어 엔지니어링, 데이터 플랫폼 및 응용 AI의 교차점에서 일하고 있습니다. 그는 특히 검색 증강 생성(RAG) 및 다중 에이전트 워크플로우와 같은 상용 및 실제 운영 수준의 생성형 AI(GenAI) 시스템 구축에 중점을 두고 있으며, 초기 아이디어를 실제 활용 가능한 환경으로 발전시키도록 팀을 돕고 있습니다. 또한 Thoughtworks의 글로벌 AI 서비스 개발 팀에 기여하고 있으며, 실제 운영 가능한 RAG 애플리케이션 구축에 대한 O'Reilly 강의를 진행하고 있습니다.
목차
과제: 임상 전 데이터 미로 탐색하기
해결책: PRINCE - 진화하는 플랫폼
시스템 아키텍처: 신뢰할 수 있는 에이전트 RAG 시스템 엔지니어링
에이전트 RAG 시스템
사용자 의도 명확히 하기
사고 및 계획: 프로세스 반성
연구원 에이전트(The Researcher Agent)
반성 에이전트(The Reflection Agent): 데이터 검증 및 충분성
작가 에이전트(The Writer Agent): 답변 종합 및 포맷팅
실제 운영 수준의 LLM 시스템에서 신뢰 구축
투명성 및 설명 가능성
평가
모니터링
회복 탄력성을 위한 엔지니어링: 오류 처리 및 복구
데이터 품질 향상: 개체명 인식 및 주석 달기
계속되는 여정: 반복적 개발
결론
임상 전(preclinical) 신약 발견은 본질적으로 복잡하고 데이터 집약적입니다. 연구원들은 이 중요한 단계에서 생성되는 방대한 양의 정보에 효율적으로 접근하고 분석해야 하는 중대한 과제에 직면해 있습니다. 일반적으로 엄격한 불리언 논리에 의존하는 전통적인 키워드 기반 검색 방법은 미묘하고 복잡한 임상 전 연구 문제 앞에서는 종종 그 한계를 드러냅니다.
대형 언어 모델(LLM)의 등장은 혁신적인 기회를 제시했습니다. LLM의 생성 능력과 정보 검색 시스템의 정밀함을 결합하면서 검색 증강 생성(RAG)이 유망한 기술로 떠올랐습니다. 이 접근 방식은 임상 전 데이터 접근 방식에 혁명을 일으킬 잠재력을 가지고 있어, 연구원들이 자연어로 복잡한 질문을 던지고 사내의 고유 데이터에 기반하여 정확하고 풍부한 문맥의 답변을 받을 수 있게 해줍니다. 바이엘은 이러한 잠재력을 일찍이 알아보고, 이 기술이 임상 전 연구의 오랜 과제를 어떻게 해결할 수 있을지 탐구하는 데 투자했습니다.
이 글에서는 우리의 여정을 공유하고자 합니다. 즉, 생성형 AI에 대한 바이엘의 초기 투자가 에이전트 RAG(Agentic RAG) 기반의 에이전트 AI 시스템인 PRINCE로 어떻게 이어졌는지 설명합니다. 이 사례 연구는 까다로운 데이터 미로였던 임상 전 데이터 검색을 직관적인 대화형 경험으로 변화시키는 과정에서의 기술적 아키텍처, 엔지니어링 결정 및 얻은 교훈을 탐구합니다. PRINCE의 수많은 엔지니어링 결정은 이제 컨텍스트 엔지니어링 및 하네스 엔지니어링이라는 관점을 통해 이해될 수 있지만, 시스템을 처음 설계할 당시에는 이러한 용어를 사용하지 않았습니다. 컨텍스트 엔지니어링은 어떤 정보를 가공하고 전달할 것인지를 형성하는 것이었습니다.
Building Reliable Agentic AI Systems A Case Study in building production-ready agentic AI systems This paper presents the Preclinical Information Center (PRINCE), a cloud-hosted platform developed by Bayer AG with Thoughtworks to address pharmaceutical industry challenges in drug development. PRINCE leverages Agentic Retrieval-Augmented Generation and Text-to-SQL to integrate decades of safety study reports. We describe PRINCE's evolution from keyword-based search to an intelligent research assistant capable of answering complex questions and drafting regulatory documents. We reflect on key engineering decisions through the lens of context engineering—how information was shaped and routed between specialized agents—and harness engineering—how orchestration, recovery, and observability were built around the models to maintain control and reliability. The system prioritizes trust through transparency, explainability, and human-in-the-loop integration. PRINCE demonstrates AI's transformative potential in pharmaceuticals, significantly improving data accessibility and research efficiency while ensuring governance and compliance. 16 June 2026 Sarang Sanjay Kulkarni Sarang Kulkarni is a Principal Consultant at Thoughtworks, working at the intersection of software engineering, data platforms, and applied AI. He focuses on building production-grade GenAI systems, particularly Retrieval-Augmented Generation (RAG) and multi-agent workflows, and helps teams take these systems from early ideas to real-world use. Sarang also contributes to Thoughtworks’ Global AI Service Development team and teaches an O’Reilly course on building production-ready RAG applications. Contents The Challenge: Navigating the Preclinical Data Maze The Solution: PRINCE - An Evolutionary Platform System Architecture: Engineering a Reliable Agentic RAG System The Agentic RAG System Clarify User Intent Think & Plan: Process Reflection The Researcher Agent The Reflection Agent: Data Validation and Sufficiency The Writer Agent: Answer Synthesis and Formatting Building Trust in a Production LLM System Transparency and Explainability Evaluation Monitoring Engineering for Resilience: Error Handling and Recovery Enhancing Data Quality: Named Entity Recognition and Annotation The Journey Continues: Iterative Development Conclusion Preclinical drug discovery is inherently complex and data-intensive. Researchers face the significant challenge of efficiently accessing and analyzing vast volumes of information generated during this critical phase. Traditional keyword-based search methods, often reliant on rigid Boolean logic, frequently fall short when confronted with the nuanced and intricate nature of preclinical research questions. The advent of Large Language Models (LLMs) has presented a transformative opportunity. By combining the generative power of LLMs with the precision of information retrieval systems, Retrieval-Augmented Generation (RAG) has emerged as a promising technique. This approach holds the potential to revolutionize preclinical data access, enabling researchers to pose complex questions in natural language and receive accurate, context-rich answers grounded in proprietary data. Recognizing this potential early, Bayer committed to exploring how these technologies could address longstanding challenges in preclinical research. In this post, we share that journey—how Bayer's early investment in generative AI has resulted in PRINCE, an agentic AI system built on Agentic RAG. This case study explores the technical architecture, engineering decisions, and lessons learned in transforming preclinical data retrieval from a challenging maze into an intuitive conversational experience. Many of the engineering decisions behind PRINCE can now be understood through the lens of context engineering and harness engineering, although when the system was first designed we did not use these terms. Context engineering shaped what information each model received, what it did not receive, and how context moved between specialized steps such as research, reflection, and writing. Harness engineering shaped the scaffolding around the models: orchestration, tool boundaries, state persistence, retries, fallbacks, validation, reflection loops, observability, and human review. While this post focuses on the technical architecture and engineering challenges, our paper published in Frontiers in Artificial Intelligence covers the product evolution and business impact in more detail. The Challenge: Navigating the Preclinical Data Maze The preclinical research landscape at Bayer, like many large pharmaceutical organizations, is characterized by a diverse and extensive array of data. This includes highly structured datasets from various studies, alongside vast amounts of unstructured information embedded within text documents such as study reports, publications, and regulatory submissions. Researchers frequently encountered significant hurdles in accessing and analyzing this information effectively: Data Silos: information was fragmented and scattered across numerous disparate systems and repositories, making it exceedingly difficult to gain a comprehensive, holistic view of preclinical data related to a specific compound or study. Limited Search Capabilities: traditional keyword-based search engines struggled with the complexity and variability of preclinical terminology and research questions, often yielding irrelevant, incomplete, or overwhelming results. Time-Consuming Manual Analysis: extracting specific insights or compiling information across multiple documents required considerable manual effort, diverting valuable researcher time away from core scientific activities. These inherent challenges highlighted a clear need for a more efficient, intelligent, and integrated approach to preclinical data retrieval and analysis. The Solution: PRINCE - An Evolutionary Platform To address these challenges, Bayer developed the Preclinical Information Center (PRINCE) platform. PRINCE was conceived as a unified gateway to preclinical data, initially focusing on consolidating previously siloed structured study metadata and exposing them in a “Searchable” manner. This initial phase allowed users to apply advanced filters and retrieve information primarily from structured study metadata. However, a significant portion of Bayer's valuable preclinical knowledge resides within unstructured PDF study reports accumulated over decades. Due to numerous system migrations over the years, the structured metadata associated with these reports could be incomplete, missing, or even contain incorrect annotations. Crucially, the authoritative “gold standard” information was consistently present within the approved PDF study reports. The emergence of Generative AI, particularly RAG, provided the key to unlocking this wealth of unstructured data. By integrating RAG capabilities, PRINCE began to shift the paradigm from a filter-based 'search' tool to a natural language 'ask' system, enabling researchers to query the content of these study reports directly. This evolution reflects PRINCE's progression through three distinct phases: Search: the initial phase focused on creating a unified gateway to thousands of nonclinical study reports, consolidating multiple in-house data silos from various preclinical domains into a searchable format, primarily leveraging structured metadata. Ask: this phase introduced an AI-powered question-answering system utilizing Retrieval Augmented Generation (RAG). This enabled researchers to derive insights directly from unstructured data, including scanned PDFs from historical reports, by posing questions in natural language. Do: the current phase positions PRINCE as an active research assistant capable of executing complex tasks. This is achieved through the integration of multi-agent systems, allowing the platform to handle intricate queries, orchestrate workflows, and sup