EC

Hire Esme C. - Prompt Engineer - San Francisco

Generalist Evaluator Expert

10+ years
san francisco, california, united states
Mercor Intelligence

About esme

Hi there! I’m an Applied AI and GenAI specialist with a background in computer science and linguistics, passionate about making language models more useful, human-centered, and reliable.My career began in tech consulting and front-end development at Accenture, where I built intuitive, user-facing web applications. Since then, I’ve pivoted into AI and content development, applying my understanding of language, logic, and UX to craft and refine prompts, design learning materials, and develop LLM-powered workflows.Currently, I’m with Mercor, leading prompt design, and rubric evaluation—writing clear, engaging guides and hands-on tutorials, testing for precision and task success, and iterating based on user feedback. I’m especially interested in how natural-language interfaces and well-designed content can empower non-technical users, and I’m building projects around content generation, information retrieval, and workflow automation.I’m seeking opportunities to join teams building AI-powered products—particularly in roles that bridge content development, AI, and engineering. Whether you’re developing internal learning platforms, customer-facing chatbots, or AI assistants, I’d love to help design better conversations between humans and machines.Areas of focus: prompt engineering, content development & automation, natural language UX, prompt evaluation, LLM product design

Key Skills

a/b testingai content evaluationai promptingai response evaluation and quality assuranceapi integrationautomationbootstrapc (programming language)c++capcutcascading style sheetschatgptcomputational linguisticscontent marketingcontent strategy+41 more

Experience

Generalist Evaluator Expert

Current

Mercor Intelligence

* Designed, tested, and refined LLM prompts, templates, and evaluation workflows to assess accuracy, reasoning quality, instruction adherence, and policy alignment across multiple domains. * Built and maintained golden evaluation sets and scoring rubrics to benchmark model behavior against quality, safety, and trust & safety standards. * Analyzed model outputs to identify failure modes, edge cases, hallucinations, and misclassifications, translating findings into prioritized improvement recommendations. * Categorized large volumes of textual and multimodal content with high precision and consistency, ensuring alignment with detailed evaluation guidelines and quality standards. * Evaluated both textual and audio outputs for accuracy, relevance, and adherence to prompt requirements, demonstrating comfort with multimodal materials. * Followed comprehensive rubric and style guides consistently, ensuring reliable decisions across thousands of examples and minimizing inter-annotator variability. * Provided clear, structured feedback on edge cases, ambiguous outputs, and guideline gaps to improve evaluation processes. * Promoted to Reviewer and Team Lead, owning quality governance processes, mentoring evaluators, and coordinating feedback loops and incident reviews to ensure consistent outcomes across teams.

Prompt Engineer

Dataannotation

* Designed and evaluated prompts for diverse NLP tasks, including classification, summarization, and customer support simulation, using LLMs like GPT-4 and Claude. * Iteratively refined instructions and prompt structures to improve model output quality, alignment, and task success rates. * Documented prompt behavior and edge cases to support scalable prompt libraries and evaluation frameworks. * Developed intuition for few-shot prompting, instruction tuning, and hallucination mitigation through hands-on experimentation. * Collaborated with internal reviewers to improve quality assurance and consistency in output evaluations.

Head Frontend Developer

Accenture

* Collaborated cross-functionally with product managers, designers, and engineers to translate complex user workflows into scalable technical implementations. * Built production web features supporting high-traffic consumer platforms, ensuring reliability, accessibility (WCAG), and consistent user experience. * Participated in iterative development cycles including QA validation, product testing, and performance monitoring. * Contributed to defining feature success metrics and evaluation strategies used to measure product impact and system reliability.

Education

University Of Illinois Urbana - Champaign

Bachelor Of Science

Girls Who Code

Leland High School

Interested in connecting with esme?

Sign up for NinjaHire to send a connection request.

Common Questions

What is esme's expertise?

esme specializes in Prompt Engineer - San Francisco, with expertise in a/b testing, ai content evaluation, ai prompting, ai response evaluation and quality assurance, api integration.

Where is esme located?

esme is based in san francisco, california, united states.

How much experience does esme have?

esme has 10+ years of professional experience.

How can I contact esme?

You can connect with esme through NinjaHire by signing up for a free account.

Looking for a different Prompt Engineer - San Francisco?

Describe exactly who you need and NinjaHire will source them for you.

Type a role to try NinjaHire for free