IA · Ciudad de México · Technology and consumer platforms
Mexican and US Hispanic Spanish training data for a United States technology group
De un vistazo
- Industry
- Technology and consumer platforms
- Client geography
- United States, serving US Hispanic and Mexican markets
- Client size
- Enterprise, consumer technology group
- Service line
- AI delivery, multilingual training data, annotation, RLHF
- Primary language
- Spanish across Mexican and US Hispanic variants, with English
- Delivery site
- Mexico City (CDMX)
- Engagement duration
- 11 months, ongoing
- Team size
- 58 annotators, 6 linguistic leads, 4 QA specialists, 2 ML liaison engineers
Perfil del cliente
The client is a United States consumer technology group whose products serve both the US Hispanic market and Mexico. Its customer-facing automation, intent classification, conversational support and content categorisation, is built on models trained predominantly on Castilian and generic Spanish data.
El reto
The client's models performed well on generic Spanish and poorly on the Spanish its users actually speak. Intent classification F1 stood at 90.5% on Castilian Spanish and fell to 76.4% on US Hispanic Spanish and 79.0% on Mexican Spanish. Those two variants were exactly the client's core market.
The failure modes were specific: vocabulary divergence, differing politeness and directness conventions, US Hispanic code-switching between Spanish and English, and Mexican idiom the generic models had never seen.
The client's initial remediation, machine-translating Castilian data, improved benchmark scores while leaving real-world performance largely unchanged, because the data was grammatically Spanish and pragmatically foreign.
Sourcing genuine Mexican and US Hispanic Spanish annotation at scale had proven difficult, because the variants required native pragmatic judgement the client could not assemble internally.
Por qué Corpshore México
Corpshore Mexico proposed Mexico City, whose scale and diversity give access to native Mexican Spanish and, through its connection to the US Hispanic experience, familiarity with US Hispanic register and code-switching. Corpshore could evidence multi-variant capability that competing single-market providers could not.
The client ran a blind trial across three providers. Corpshore's inter-annotator agreement on Mexican and US Hispanic Spanish was materially higher, and its linguistic leads produced a written taxonomy of the pragmatic divergences causing misclassification. Corpshore also proposed embedding ML liaison engineers in the client's pipeline rather than delivering data as a batch handoff.
La colaboración
Fifty-eight annotators in Mexico City organised into variant desks, Mexican, US Hispanic, Central American, Caribbean and Castilian, with six linguistic leads, four QA specialists and two ML liaison engineers working inside the client's pipeline. Annotators are native speakers of the variant they annotate. All complete a four-week onboarding covering the label schema, pragmatic annotation principles and the divergence taxonomy.
Enfoque y metodología
Native variant annotation, never translation. Every record is annotated by a native speaker of its variant, using data originally produced in that variant.
Pragmatic annotation, not only semantic. The label schema captures directness, politeness register, urgency, code-switching and sentiment intensity, the features that vary and cause misclassification.
Active learning against model uncertainty. Annotation priority is set weekly by the model's uncertainty, reducing total annotation volume required against uniform sampling.
Cross-variant calibration. All desks review a shared weekly sample to establish where a concept is the same across variants and where it differs, which requires the desks to be co-located.
Resultados
Intent classification F1 rose above 91% across all variants, with US Hispanic Spanish improving from 76.4% to 92.9% and Mexican from 79.0% to 92.6%. The spread between best and worst variant narrowed from 18.9 points to 1.7.
The client launched conversational automation for its US Hispanic and Mexican audiences on schedule, having previously deferred those launches on model performance grounds.
Annotation throughput reached 63,400 records per week by month six. Cost per annotated record was 58% below the client's prior vendor benchmark.
Indicadores clave
| Indicador | Inicial | Después | Cambio |
|---|---|---|---|
| Intent F1, US Hispanic Spanish | 76.4% | 92.9% | +16.5 pts |
| Intent F1, Mexican Spanish | 79.0% | 92.6% | +13.6 pts |
| Intent F1, Central American Spanish | 78.1% | 92.2% | +14.1 pts |
| Intent F1, Caribbean Spanish | 71.6% | 91.7% | +20.1 pts |
| Spread across variants | 18.9 pts | 1.7 pts | -91% |
| Weekly annotation throughput | n/a | 63,400 records | New capability |
| Cost per annotated record | Prior vendor baseline | -58% | -58% |
We had treated Spanish as one thing. The divergence taxonomy their linguistic leads produced in the trial changed how our whole ML team thinks about the US Hispanic and Mexican markets.
Valor duradero
The pragmatic divergence taxonomy and extended label schema are client-owned and govern all the client's Spanish-language model work. The multi-variant single-site model has become a defined Corpshore Mexico capability. The engagement has extended into RLHF for the client's generative assistant.
Temas
Corpshore Mexico