Computational Linguist · AI Governance & Multilingual NLP Researcher
MS Computer Science (Incoming, Georgia Institute of Technology) | BA Applied Linguistics (University of California, Los Angeles)

Sahara Al-Madi

“How do we govern AI systems that serve multilingual communities?” Founder of the Linguistic Security Institute (LSI). Specializing in leading multilingual AI data operations, LLM safety guardrail alignment, dialectal bias auditing, and cross-lingual security protocols.

Professional Experience

Google (via TEKsystems) — AI Language Consultant

May 2026–Present

Multilingual AI Golden Label Sets
Data Operations
Localization
LLM Evaluation

SoundHound AI — AI Data Operations Consultant

Oct 2024–May 2026

Voice AI
QA Frameworks
Data Annotation
Project Leadership

Linguistic Security Institute (LSI) — Founder & Lead Researcher

Sep 2025–Present

The Linguistic Security Institute (LSI) is an open-source research initiative dedicated to ensuring AI systems are culturally coherent, institutionally resilient, and ethically verifiable. We audit linguistic bias in AI, develop open-source tools for bias detection, create educational resources for developers and students, and propose governance frameworks for linguistic safety in AI development. Visit LSI →

AI Governance
Bias Auditing
Open-Source Research

National Museum of Language — Digital Strategy Lead

Jan 2020–Aug 2025

Directed multilingual digital strategy, expanding organic web traffic by 62% and overseeing cross-functional transcription/translation programs in English, Spanish, and Arabic.

Multilingual Strategy
Grant Writing
Cross-Functional Leadership

Core Expertise & Technical Stack

Programming & Tooling

Python (Pandas, NumPy), SQL, Git, GitHub, Bash environment scripting.

AI & NLP Methodologies

LLM Evaluation, Error Analysis, RAG architectures, Prompt Engineering, Zero-Shot Learning, Bias Analysis.

Data & Quality Assurance

Data Annotation, Localization workflows, Gold Label Audits, Rigorous Dataset Validation.

Linguistics & Languages

Dialectology, IPA/SAMPA.
English & Spanish (Native), Arabic (Writing & Linguistic Analysis: Working Proficiency)  Italian (Intermediate).

Research & Publications

ACL 2026 · SemEval Task 9

Arabic Polarization Detection Across 10+ Large Language Models

Second-author and Error Analysis Lead. Investigated dialectal bias, hallucinations, and Clever Hans effects across 10+ models (DeepSeek, Qwen, Gemma, Llama, Fanar). View Paper →

Coming Soon

New Announcement Coming Soon

Coming Soon

New Announcement Coming Soon

Open Source Projects & Workshops

Linguistic Firewall: Polyglot Poisoning

Open-source Python toolkit for detecting multilingual prompt injection and polyglot poisoning vulnerabilities.  An expanded, reproducible version of the demo originally presented in April 2026 will be featured as an official hands-on workshop at DEATHCon BSides on November 14, 2026.

Mitote

RAG text-to-speech prototype supporting Nahuatl-aware Spanish pronunciation for Indigenous language preservation.

2026 Speaking Engagements & Poster Showcase

Official LSI 2026 Engagements & Research Poster Showcase

© 2026 Sahara Al-Madi · Linguistic Security Institute, LLC