MS Computer Science (Incoming, Georgia Institute of Technology) | BA Applied Linguistics (University of California, Los Angeles)
Sahara Al-Madi
“How do we govern AI systems that serve multilingual communities?” Founder of the Linguistic Security Institute (LSI). Specializing in leading multilingual AI data operations, LLM safety guardrail alignment, dialectal bias auditing, and cross-lingual security protocols.
Professional Experience
Google (via TEKsystems) — AI Language Consultant
May 2026–Present
Data Operations
Localization
LLM Evaluation
SoundHound AI — AI Data Operations Consultant
Oct 2024–May 2026
QA Frameworks
Data Annotation
Project Leadership
Linguistic Security Institute (LSI) — Founder & Lead Researcher
Sep 2025–Present
The Linguistic Security Institute (LSI) is an open-source research initiative dedicated to ensuring AI systems are culturally coherent, institutionally resilient, and ethically verifiable. We audit linguistic bias in AI, develop open-source tools for bias detection, create educational resources for developers and students, and propose governance frameworks for linguistic safety in AI development. Visit LSI →
Bias Auditing
Open-Source Research
National Museum of Language — Digital Strategy Lead
Jan 2020–Aug 2025
Directed multilingual digital strategy, expanding organic web traffic by 62% and overseeing cross-functional transcription/translation programs in English, Spanish, and Arabic.
Grant Writing
Cross-Functional Leadership
Core Expertise & Technical Stack
Programming & Tooling
Python (Pandas, NumPy), SQL, Git, GitHub, Bash environment scripting.
AI & NLP Methodologies
LLM Evaluation, Error Analysis, RAG architectures, Prompt Engineering, Zero-Shot Learning, Bias Analysis.
Data & Quality Assurance
Data Annotation, Localization workflows, Gold Label Audits, Rigorous Dataset Validation.
Linguistics & Languages
Dialectology, IPA/SAMPA.
English & Spanish (Native), Arabic (Writing & Linguistic Analysis: Working Proficiency) Italian (Intermediate).
Research & Publications
Arabic Polarization Detection Across 10+ Large Language Models
Second-author and Error Analysis Lead. Investigated dialectal bias, hallucinations, and Clever Hans effects across 10+ models (DeepSeek, Qwen, Gemma, Llama, Fanar). View Paper →
New Announcement Coming Soon
New Announcement Coming Soon
Open Source Projects & Workshops
Linguistic Firewall: Polyglot Poisoning
Open-source Python toolkit for detecting multilingual prompt injection and polyglot poisoning vulnerabilities. An expanded, reproducible version of the demo originally presented in April 2026 will be featured as an official hands-on workshop at DEATHCon BSides on November 14, 2026.
Mitote
RAG text-to-speech prototype supporting Nahuatl-aware Spanish pronunciation for Indigenous language preservation.
2026 Speaking Engagements & Poster Showcase
Official LSI 2026 Engagements & Research Poster Showcase
