CAS uses unparalleled scientific content, specialized technology and unmatched human expertise to help R&D organizations across Commercial, Government and Academic sectors create groundbreaking innovations that benefit the world. As the Scientific Information Solutions Division of the American Chemical Society, CAS manages the largest curated reservoir of scientific knowledge, and for 119 years, has helped innovators mine, assess and apply that information to keep businesses thriving. The CAS team is global, diverse, endlessly curious and strives to make actionable scientific insights accessible to innovators worldwide.
CAS is currently seeking a Senior Data Scientist focused on agentic AI development and large language model (LLM) engineering. This position will be located at our headquarters in Columbus, Ohio.
This role sits on the Data Analytics and Insights (DAI) team, which builds the AI that powers CAS's scientific information products. You will work on the systems behind the Newton research assistants in SciFinder and BioFinder, the natural-language query agent in IPFinder, predictive models that ship inside products, and AI-assisted content curation pipelines, alongside data engineers, product managers, and scientific domain experts.
This role requires a highly self-directed professional who can independently design, build, and evaluate agentic AI systems and LLM-powered applications that turn CAS's curated, provenanced scientific content into trustworthy research capabilities. The successful candidate will provide strategic input into product and platform ideation, drive complex technical initiatives from concept through production, and translate advanced AI engineering into tangible business value with minimal oversight, while providing technical direction to less-experienced scientists and engineers.
Key Responsibilities
Agentic AI & LLM System Development
- Independently design, build, and productionize agentic AI systems — including multi-step, tool-using, and multi-agent workflows — that support scientific research tasks.
- Develop retrieval-augmented generation (RAG) pipelines that ground model outputs in curated, provenanced scientific content.
- Implement and extend model-context and tool-integration frameworks (for example, the Model Context Protocol) to connect LLMs and agents to scientific data sources and external systems.
- Apply prompt engineering, context engineering, and orchestration techniques to improve the reliability and quality of agentic and generative workflows.
Applied AI Engineering & Model Adaptation
- Fine-tune, adapt, and deploy large language models — including self-hosted, open-weight models — for domain-specific scientific tasks.
- Develop and deploy machine learning models that ship inside CAS products, such as property, toxicity, or biologic developability prediction, using scientific judgment to inform feature design and validate model behavior.
- Take AI capabilities from prototype all the way through production: building the data, training, and inference pipelines, owning the deployment, and iterating after first release, not just the proof of concept.
- Work with engineering and product to put AI features in front of real users at scale, and stay accountable for how they behave once they're live.
- Set and hold the engineering standard for production AI: appropriate safeguards such as content-safety guardrails, entitlement-based access control, evaluation, and model fallback, delivered as clean, tested, well-documented software with unit, integration, and end-to-end tests, containerized development, and CI/CD.
- Use agentic software development tools such as Claude Code as a core part of the daily workflow, and set the team's practice for using them well.
Scientific Grounding & Evaluation
- Design rigorous evaluation frameworks for LLM and agentic systems — measuring accuracy, faithfulness, and robustness, and distinguishing acceptable model variability from factual error.
- Critically evaluate scientific datasets and model outputs for quality, completeness, and scientific validity, applying domain understanding of chemistry, materials science, or the life sciences.
- Establish trustworthy-AI practices that keep model outputs grounded in curated, provenanced scientific content rather than the model's own priors.
Business Impact & Communication
- Independently synthesize technical findings and present strategic recommendations to senior executives and C-level stakeholders.
- Influence organizational and product decisions through compelling, data-driven narratives and demonstrations.
- Represent CAS's AI work externally through client engagements, conference presentations, and publications.
Technical Leadership & Mentorship
- Set the technical direction for a team of data scientists and data engineers: the standards, patterns, and architecture for how agentic and LLM systems get built, spanning data and inference pipelines, model and prompt design, and deployment.
- Own design and code review across that work, from model and prompt design through data pipelines and serving infrastructure, and raise the bar on quality, reliability, and reproducibility.
- Mentor junior data scientists and AI/data engineers, and unblock them on difficult technical problems.
- Set the team's standards for AI engineering, evaluation, and responsible deployment.
- Lead and influence cross-functional initiatives as part of a shared-services model, without direct people-management responsibility.
- Share knowledge across the team through demos, documentation, and internal learning sessions.
Qualifications
Required
- Master's degree in Computer Science, Data Science, Chemistry, Materials Science, Chemical Engineering, or a related technical or scientific field with 6+ years of applied AI/ML or data science experience; a Bachelor's degree in one of those fields with 8+ years; or a PhD with 3+ years.
- Advanced proficiency in Python and the modern AI engineering stack (for example, PyTorch, Hugging Face, vector databases, orchestration/agent frameworks, SQL, and cloud platforms).
- Demonstrated, hands-on experience building LLM-powered applications, agentic or multi-agent systems, and/or RAG pipelines — and taking them to real users in production, including the post-launch iteration that implies.
- Experience fine-tuning, adapting, or deploying large language models, including familiarity with self-hosted or open-weight models.
- Working knowledge of evaluation methodology for generative and agentic systems (accuracy, faithfulness, and hallucination/variability management).
- Ability to work with scientific data and domain concepts — such as chemical or materials data, molecular representations, or life-sciences datasets — well enough to judge whether a system's outputs are accurate, defensible, and useful to working scientists, and to partner effectively with subject-matter experts.
- Exceptional communication skills, with a proven ability to influence senior stakeholders and lead cross-functional efforts.
- Strong project-management capability across multiple concurrent initiatives.
- Experience deploying and operating production systems in at least one cloud environment (AWS, Azure, or GCP).
- Proficiency with agentic software development tools (for example, Claude Code or similar AI-assisted development environments).
Preferred
- PhD in a computational, physical, chemical, or life-sciences field.
- Experience with tool-use and integration frameworks such as the Model Context Protocol (MCP).
- Experience self-hosting or serving open-weight models in production (LLMOps / MLOps).
- Cheminformatics experience or familiarity with molecular representations and scientific databases.
- Experience building and deploying predictive ML models in a scientific or regulated domain.
- Track record of publications or conference presentations in AI-for-science or applied machine learning.
- Consulting or client-facing experience in scientific, chemical, or life-sciences sectors.
CAS offers a competitive salary and comprehensive benefits package, including a generous vacation plan, medical, dental, vision insurance plans, and employee savings and retirement plans. Candidates for this position must be authorized to work in the United States and not require work authorization sponsorship by our company for this position now or in the future. EEO/Minority/Female/Disabled/Veteran
Equal Opportunity Employer/Protected Veterans/Individuals with DisabilitiesThis employer is required to notify all applicants of their rights pursuant to federal employment laws. For further information, please review the Know Your Rights notice from the Department of Labor.
Skills Required
- Master's degree in Computer Science, Data Science, Chemistry, Materials Science, Chemical Engineering, or related technical/scientific field with 6+ years of applied AI/ML or data science experience
- Bachelor's degree in a relevant field with 8+ years of applied AI/ML or data science experience, or PhD with 3+ years
- Advanced Python proficiency and experience with modern AI engineering technologies including PyTorch, Hugging Face, vector databases, agent frameworks, SQL, and cloud platforms
- Hands-on experience building and productionizing LLM-powered applications, agentic or multi-agent systems, and/or RAG pipelines
- Experience fine-tuning, adapting, or deploying large language models, including self-hosted or open-weight models
- Working knowledge of generative and agentic AI evaluation methodologies, including accuracy, faithfulness, and hallucination management
- Ability to work with scientific data and domain concepts including chemical, materials, molecular, or life-sciences data
- Exceptional communication skills and ability to influence senior stakeholders and lead cross-functional efforts
- Strong project-management capability across multiple concurrent initiatives
- Experience deploying and operating production systems in AWS, Azure, or GCP
- Proficiency with agentic software development tools such as Claude Code or similar AI-assisted development environments
- PhD in a computational, physical, chemical, or life-sciences field
- Experience with tool-use and integration frameworks such as Model Context Protocol
- Experience self-hosting or serving open-weight models in production, including LLMOps or MLOps
- Cheminformatics experience or familiarity with molecular representations and scientific databases
- Experience building and deploying predictive ML models in a scientific or regulated domain
- Publications or conference presentations in AI-for-science or applied machine learning
- Consulting or client-facing experience in scientific, chemical, or life-sciences sectors
What We Do
At CAS, we curate, connect, and analyze scientific knowledge to reveal the unseen connections that inspire breakthroughs. We weave a fabric of discovery that scientific innovators can tap into to stimulate their creativity and accelerate their work. Because when the world turns to science, science turns to CAS. So if you're advancing research, repurposing technology, making strategic decisions, or leading digital R&D initiatives, take a look at our story below to see how partnering with us gets you there faster. A global company based in Columbus, Ohio, CAS employs over 1,400 experts who curate, connect, and analyze scientific knowledge to reveal unseen connections. CAS is a division of the American Chemical Society. Connect with us at cas.org.








