Skip to content
Corpshore Mexico

AI · Mexico City · Technology and consumer platforms

Mexican and US Hispanic Spanish training data for a United States technology group

At a glance

Industry
Technology and consumer platforms
Client geography
United States, serving US Hispanic and Mexican markets
Client size
Enterprise, consumer technology group
Service line
AI delivery, multilingual training data, annotation, RLHF
Primary language
Spanish across Mexican and US Hispanic variants, with English
Delivery site
Mexico City (CDMX)
Engagement duration
11 months, ongoing
Team size
58 annotators, 6 linguistic leads, 4 QA specialists, 2 ML liaison engineers

Client profile

The client is a United States consumer technology group whose products serve both the US Hispanic market and Mexico. Its customer-facing automation, intent classification, conversational support and content categorisation, is built on models trained predominantly on Castilian and generic Spanish data.

The challenge

The client's models performed well on generic Spanish and poorly on the Spanish its users actually speak. Intent classification F1 stood at 90.5% on Castilian Spanish and fell to 76.4% on US Hispanic Spanish and 79.0% on Mexican Spanish. Those two variants were exactly the client's core market.

The failure modes were specific: vocabulary divergence, differing politeness and directness conventions, US Hispanic code-switching between Spanish and English, and Mexican idiom the generic models had never seen.

The client's initial remediation, machine-translating Castilian data, improved benchmark scores while leaving real-world performance largely unchanged, because the data was grammatically Spanish and pragmatically foreign.

Sourcing genuine Mexican and US Hispanic Spanish annotation at scale had proven difficult, because the variants required native pragmatic judgement the client could not assemble internally.

Why Corpshore Mexico

Corpshore Mexico proposed Mexico City, whose scale and diversity give access to native Mexican Spanish and, through its connection to the US Hispanic experience, familiarity with US Hispanic register and code-switching. Corpshore could evidence multi-variant capability that competing single-market providers could not.

The client ran a blind trial across three providers. Corpshore's inter-annotator agreement on Mexican and US Hispanic Spanish was materially higher, and its linguistic leads produced a written taxonomy of the pragmatic divergences causing misclassification. Corpshore also proposed embedding ML liaison engineers in the client's pipeline rather than delivering data as a batch handoff.

The engagement

Fifty-eight annotators in Mexico City organised into variant desks, Mexican, US Hispanic, Central American, Caribbean and Castilian, with six linguistic leads, four QA specialists and two ML liaison engineers working inside the client's pipeline. Annotators are native speakers of the variant they annotate. All complete a four-week onboarding covering the label schema, pragmatic annotation principles and the divergence taxonomy.

Approach and methodology

Native variant annotation, never translation. Every record is annotated by a native speaker of its variant, using data originally produced in that variant.

Pragmatic annotation, not only semantic. The label schema captures directness, politeness register, urgency, code-switching and sentiment intensity, the features that vary and cause misclassification.

Active learning against model uncertainty. Annotation priority is set weekly by the model's uncertainty, reducing total annotation volume required against uniform sampling.

Cross-variant calibration. All desks review a shared weekly sample to establish where a concept is the same across variants and where it differs, which requires the desks to be co-located.

Results

Intent classification F1 rose above 91% across all variants, with US Hispanic Spanish improving from 76.4% to 92.9% and Mexican from 79.0% to 92.6%. The spread between best and worst variant narrowed from 18.9 points to 1.7.

The client launched conversational automation for its US Hispanic and Mexican audiences on schedule, having previously deferred those launches on model performance grounds.

Annotation throughput reached 63,400 records per week by month six. Cost per annotated record was 58% below the client's prior vendor benchmark.

Key indicators

MetricBaselineAfterChange
Intent F1, US Hispanic Spanish76.4%92.9%+16.5 pts
Intent F1, Mexican Spanish79.0%92.6%+13.6 pts
Intent F1, Central American Spanish78.1%92.2%+14.1 pts
Intent F1, Caribbean Spanish71.6%91.7%+20.1 pts
Spread across variants18.9 pts1.7 pts-91%
Weekly annotation throughputn/a63,400 recordsNew capability
Cost per annotated recordPrior vendor baseline-58%-58%
We had treated Spanish as one thing. The divergence taxonomy their linguistic leads produced in the trial changed how our whole ML team thinks about the US Hispanic and Mexican markets.
Head of Machine Learning, US technology group

Enduring value

The pragmatic divergence taxonomy and extended label schema are client-owned and govern all the client's Spanish-language model work. The multi-variant single-site model has become a defined Corpshore Mexico capability. The engagement has extended into RLHF for the client's generative assistant.

Topics

Spanish AI training data Mexicodata annotation MexicoUS Hispanic Spanish NLPmultilingual data annotation Mexico

Corpshore Mexico