AI Job Radar

Model Evaluation Jobs – Page 4

Aktuelle KI-Jobs mit Model Evaluation, passende Lernpfade und Bewerbungsbezug.

How to use Model Evaluation in applications

If a job requires Model Evaluation, the skill should be supported by a project, course or portfolio example. The application check reviews whether the skill is actually evidenced in your CV.

67
Results
28
Companies
94.2
Average score
20
Remote

10 results on this page. 67 results in total. More results are available via pagination, company pages, skill pages and job detail pages.

PhD GenAI Research Scientist Intern

Databricks · Mountain View, California

95/100
Mountain View, CaliforniaUSAgreenhouse2026-07-31

Why this is a real AI job: The role is explicitly focused on research and development of LLMs and AI systems for enterprise applications. The job description details tasks directly related to adapting, improving, and evaluating these models.

Company Description: At Databricks, we are obsessed with enabling data teams to solve the world’s toughest problems, from security threat detection to cancer drug development. We do this by building and running the world’s best data and AI platform, so our customers can focus on the high value chal…

Details Open source / apply

Senior Staff Applied AI Engineer - Context Retrieval

Databricks · Mountain View, California; San Francisco, California

95/100
Mountain View, California; San Francisco, CaliforniaUSAgreenhouse2026-07-31

Why this is a real AI job: The role is entirely focused on building and improving context retrieval systems for AI agents, encompassing all core AI/ML/LLM aspects. The job description explicitly mentions building the retrieval stack, search subagents, and optimizing for LLMs, indicatin…

P-1549 At Databricks, we are passionate about enabling data teams to solve the world's toughest problems — from making the next mode of transportation a reality to accelerating the development of medical breakthroughs. We do this by building and running the world's best data and AI infrastructure p…

Details Open source / apply

Research Engineer, Model Evaluations

Anthropic · San Francisco, CA

95/100
San Francisco, CAUSAgreenhouse2026-07-31

Why this is a real AI job: The role is explicitly focused on evaluating and improving large language models (Claude), building evaluation infrastructure, and monitoring model health. The core responsibilities directly involve AI/ML concepts and techniques.

About Anthropic Anthropic’s mission is to create reliable, interpretable, and steerable AI systems. We want AI to be safe and beneficial for our users and for society as a whole. Our team is a quickly growing group of committed researchers, engineers, policy experts, and business leaders working to…

Details Open source / apply

Data Scientist, GTM Intelligence

OpenAI · San Francisco

95/100
San FranciscoUSAFullTimeashby2026-07-31

Why this is a real AI job: The role explicitly focuses on building and operating 'decision data products' leveraging ML models, statistical methods, and feature engineering to improve GTM strategies. The job description emphasizes productionizing these systems and monitoring their perf…

About the Team The GTM Intelligence Solutions team builds the data and decision systems that help customer-facing teams take the right action at the right time. We combine product telemetry, commercial data, customer context, and field activity to identify account health and opportunity, recommend…

Details Open source / apply

Data Scientist, Preparedness

OpenAI · San Francisco

95/100
San FranciscoUSAFullTimeashby2026-07-31

Why this is a real AI job: The role is explicitly focused on building, evaluating, and improving mitigations for harms from AI systems. It requires deep technical skills in data science, machine learning, and model evaluation, with a strong emphasis on identifying and addressing risks…

About the Team The Preparedness team is an important part of the Safety Systems https://openai.com/safety/safety-systems org at OpenAI, and is guided by OpenAI’s Preparedness Framework https://openai.com/index/updating-our-preparedness-framework/. Frontier AI models have the potential to benefit al…

Details Open source / apply

San FranciscoUSAFullTimeashby2026-07-31

Why this is a real AI job: The role is explicitly focused on building and improving AI agents (Codex), working directly with LLMs, and turning research into production systems. The tasks described are overwhelmingly centered around AI/ML concepts.

About the Team The Codex Core Agent team builds the kernel of Codex. We own making the agent better, accelerating research, and making those improvements real in production for our users. That means working across the systems that make Codex actually function as an agent in the real world: the prod…

Details Open source / apply

Data Scientist, Safety

OpenAI · San Francisco

95/100
San FranciscoUSAFullTimeashby2026-07-31

Why this is a real AI job: The role is explicitly focused on building analytical foundations for responsible AI deployment, evaluating and improving safety classifiers, and quantifying safety risks. The job description heavily emphasizes AI/ML-related tasks and requires expertise in ar…

ABOUT THE TEAM OpenAI’s Safety teams work to ensure our products are safe, trusted, and resilient as frontier AI systems scale globally. We tackle some of the company’s most important challenges across understanding and preventing misuse and misalignment, intercepting fraud and abuse, and protectin…

Details Open source / apply

Remote - United StatesUSAgreenhouse2026-07-28

Why this is a real AI job: The role is explicitly focused on building and leading machine learning systems for content understanding within the Ads organization. The responsibilities heavily emphasize ML modeling, pipelines, and deployment, with a strong focus on LLMs and NLP technique…

Reddit is a community of communities. It’s built on shared interests, passion, and trust, and is home to the most open and authentic conversations on the internet. Every day, Reddit users submit, vote, and comment on the topics they care most about. With 100,000+ active communities and approximatel…

Details Open source / apply

Machine Learning Engineer, Enterprise Brain

Glean · Mountain View, CA (HQ)

95/100
Mountain View, CA (HQ)USAgreenhouse2026-07-28

Why this is a real AI job: The role explicitly focuses on building and improving core ML models (LLMs, ranking, prediction) for a proactive AI platform. The tasks are heavily centered around AI/ML techniques and their application to enterprise workflows.

About Glean: Glean is the Work AI platform that helps everyone work smarter with AI. What began as the industry’s most advanced enterprise search has evolved into a full-scale Work AI ecosystem, powering intelligent Search, an AI Assistant, and scalable AI agents on one secure, open platform. With…

Details Open source / apply

Senior AI Engineer - AML

Adyen · Amsterdam

95/100
AmsterdamGermanygreenhouse2026-07-24

Why this is a real AI job: The role explicitly focuses on building and deploying AI systems (agentic systems, LLMs) for a core business function (AML). The description details tasks like designing agent architectures, building platform infrastructure, and evaluating model performance –…

This is Adyen Adyen provides payments, data, and financial products in a single solution for customers like Meta, Uber, H&M, and Microsoft - making us the financial technology platform of choice. At Adyen, everything we do is engineered for ambition. For our teams, we create an environment with opp…

Details Open source / apply