La IA propia de CareCompileCareCompile's own AI/ v0, en desarrollo/ v0, in development
Dicte, reúnase y encuentre lo que se dijo.Dictate, meet, and find what was said.
Vocero convierte una reunión de gobierno en un acta con decisiones y responsables, que una persona revisa. Nuestro objetivo es que lo haga en el español de El Salvador, en equipos que controlamos.
Vocero turns a government meeting into minutes with decisions and owners, which a person reviews. Our goal is for it to do so in the Spanish of El Salvador, on equipment we control.
Pedimos una sola cosa: que El Salvador designe una contraparte institucional y abramos una conversación sobre voces y textos, bajo el marco ético.
We ask for one thing: that El Salvador designate an institutional counterpart and we open a conversation about voices and texts, under the ethics framework.
Stage: in development, internal use. Our own model is not trained yet. Demos are synthetic.
PVista de presentación: muestra solo lo esencial. Pulse P otra vez para salir.Presentation view: shows only the essentials. Press P again to exit.
Synthetic exampleSo we close the budget in this phase.
Illustrative animation. It does not use your microphone and does not show model output.
01 / Demostración: reunión01 / Demo: meeting
From conversation to minutes.
A synthetic meeting, processed step by step: the ear recognizes each turn, the brain extracts entities, decisions and tasks, and every point in the minutes cites its source line. Hover over a point in the minutes to see where it comes from.
Synthetic: sample script, not measured performance.
Audiomeeting signal
Earspeech recognition
Textspeaker, time and words
Brainlanguage model: decisions, tasks, entities
Minuteseach point with its source line
Listening00:00 / 00:24
Marta0 s
Carlos0 s
Elena0 s
1
Marta00:00
Westartwiththetrainingplanforthenorthernregions.
Placenorthern regionsExtracted: place, region: north
PersonCarlosExtracted: person, owner of the sitesDatethis weekExtracted: date, deadline: this weekPersonElenaExtracted: person, owner of the requestDateThursdayExtracted: date, deadline: Thursday
5
Carlos00:20
Done,andIwillflagitiftheequipmentislate.
Minutes
The minutes are assembled when the brain finishes reading the transcript.
Decisions
Phase one is covered by the current budget.L3
Phase two is requested before quarter close.L3
Tasks
Confirm the sites
Owner
Carlos
Due
this week
L4
Prepare the phase two request
Owner
Elena
Due
Thursday
L3,4
Open
The fourth site depends on the computer equipment arriving.L2,5
{
"synthetic": true,
"meeting": "Planning call, ministry",
"decisions": [
{
"text": "Phase one is covered by the current budget.",
"source_lines": [
3
]
},
{
"text": "Phase two is requested before quarter close.",
"source_lines": [
3
]
}
],
"tasks": [
{
"task": "Confirm the sites",
"owner": "Carlos",
"due": "this week",
"source_lines": [
4
]
},
{
"task": "Prepare the phase two request",
"owner": "Elena",
"due": "Thursday",
"source_lines": [
3,
4
]
}
],
"open_items": [
{
"text": "The fourth site depends on the computer equipment arriving.",
"source_lines": [
2,
5
]
}
],
"entities": [
{
"type": "place",
"text": "northern regions",
"value": "region: north",
"line": 1
},
{
"type": "amount",
"text": "three sites",
"value": "3 sites ready, 1 pending",
"line": 2
},
{
"type": "date",
"text": "quarter close",
"value": "deadline: quarter close",
"line": 3
},
{
"type": "person",
"text": "Carlos",
"value": "owner of the sites",
"line": 4
},
{
"type": "date",
"text": "this week",
"value": "deadline: this week",
"line": 4
},
{
"type": "person",
"text": "Elena",
"value": "owner of the request",
"line": 4
},
{
"type": "date",
"text": "Thursday",
"value": "deadline: Thursday",
"line": 4
}
]
}
Type a question or pick a suggestion. Every answer cites the lines it comes from; if the meeting does not say it, Vocero says so instead of inventing.
What it extracted0 / 7
Person0
They appear as the audio is recognized.
Date0
They appear as the audio is recognized.
Amount0
They appear as the audio is recognized.
Place0
They appear as the audio is recognized.
Everything in this demo is a synthetic script running in your browser: timings, waveforms and entities illustrate the output format, not model performance.
Dictado con las reglas realesDictation with the real rules
Try it
Type what a person would say aloud, with the spoken commands: "coma", "punto", "nueva línea", "no perdón". You will see how the dictation rules turn it into written text.
Deterministic rule layer. Everything happens in your browser: no microphone, nothing is sent.
estimados miembros del comité coma el presupuesto es de cuarenta mil punto nueva línea gracias punto
Written text
Estimados miembros del comité, el presupuesto es de 40,000.
Gracias.
Commands applied
coma becomes →,
cuarenta mil becomes →40,000
punto becomes →.
nueva línea becomes →new line
punto becomes →.
Flagged ambiguities
No ambiguity flagged in this text.
Kept as words
No command was kept as a word.
This shows only the dictation rules; speech recognition is not trained yet.
02 / Datos para El Salvador02 / Data for El Salvador
Ayúdenos a que Vocero entienda cómo se habla en El Salvador.Help Vocero understand how El Salvador speaks.
Buscamos voces y textos de todo el país, aportados con consentimiento y bajo reglas claras. Lo que se aporte sirve solo para construir y evaluar Vocero.
We are looking for voices and texts from across the country, given with consent and under clear rules. What is contributed is used only to build and evaluate Vocero.
Buscamos voces de los 14 departamentos.We are looking for voices from all 14 departments.Hoy no hay datos de ningún departamento: se muestran los límites, no cobertura.There is no data from any department today: the map shows boundaries, not coverage.
Límite de un departamentoDepartment boundaryDepartamento elegidoSelected departmentGeometría: Natural Earth, dominio público.Geometry: Natural Earth, public domain.
Qué necesitamos
What we need
Audio leído con guion y habla espontánea, con su transcripción.
Voces de todas las regiones, edades y géneros.
Textos públicos con licencia clara.
Revisores nativos para calificar actas.
Scripted and spontaneous speech, with transcripts.
Voices from every region, age and gender.
Public texts with a clear license.
Native reviewers to score minutes.
Quién puede aportar
Who can contribute
Instituciones públicas, con textos de dominio público.
Universidades: estudiantes y docentes que quieran grabar.
Medios de comunicación, con archivos que puedan licenciar.
Personas voluntarias mayores de edad.
Public institutions, with public-domain texts.
Universities: students and staff who want to record.
Media outlets, with archives they can license.
Adult volunteers.
Cómo se gobierna
How it is governed
Consentimiento escrito e informado.
Propósito limitado: solo construir y evaluar Vocero.
Derecho a retirarse y a que se borren sus datos.
Los datos se guardan en equipos de CareCompile. Alojarlos dentro de El Salvador es una propuesta, aún no decidida.
Nada se revende.
Written, informed consent.
Limited purpose: only to build and evaluate Vocero.
The right to withdraw and have the data deleted.
Data is kept on CareCompile equipment. Hosting it inside El Salvador is a proposal, not yet decided.
Nothing is resold.
Qué no aceptamos
What we do not accept
Datos reales de pacientes.
Conversaciones privadas de cualquier persona sin su consentimiento.
Grabaciones de menores de edad.
Material sin derecho a compartirlo.
Real patient data.
Anyone's private conversations without their consent.
Recordings of minors.
Material you have no right to share.
¿Su institución, universidad o medio quiere participar?Would your institution, university or outlet like to take part?
Escríbanos por el formulario. Le explicamos el consentimiento y el proceso antes de que se grabe nada.Write to us through the form. We explain the consent and the process before anything is recorded.
03 / Por qué será difícil de copiar03 / Why this will be hard to copy
Lo que estamos construyendo, y lo que es verdad hoy.What we are building, and what is true today.
Hoy Vocero no tiene un foso. Tiene un plan para construirlo y herramientas que lo hacen creíble. Nada está entrenado y todavía no hay clientes.
Today Vocero has no moat. It has a plan to build one and tools that make the plan credible. Nothing is trained and there are no customers yet.
01En preparaciónIn preparation
Habla regional con consentimientoConsented regional speech
Por qué cuesta copiarloWhy it is hard to copy
Es habla de hablantes de la región, con consentimiento para cada uso y derecho a retirarlo. No se puede descargar ni comprar, solo se gana con confianza.It is speech from speakers across the region, with consent for each use and a right to withdraw it. It cannot be downloaded or bought, only earned through trust.
Qué existe hoyWhat exists today
Un borrador del consentimiento y un plan de 90 días. Ninguna hora recolectada.A consent draft and a 90-day plan. No hours collected.
Qué lo erosionaríaWhat would erode it
Un modelo abierto que añada la región, o un consentimiento defectuoso, que dejaría sin valor los datos.An open model that adds the region, or defective consent, which would make the data worthless.
02DiseñadoDesigned
Examen sellado por paísSealed exam per country
Por qué cuesta copiarloWhy it is hard to copy
Es un examen por país guardado por alguien que no entrena modelos, de modo que nadie pueda ajustar el modelo al examen.It is one exam per country, held by someone who does not train models, so nobody can tune a model to the exam.
Qué existe hoyWhat exists today
El diseño y el método, escritos. Todavía no hay un examen sellado.The design and the method, written down. There is no sealed exam yet.
Qué lo erosionaríaWhat would erode it
Que el examen se filtre, que parezca interesado si no lo custodia alguien ajeno, o que otro examen público sea mejor.The exam leaking, looking self-serving without an outside custodian, or a better public exam appearing.
03Buscamos contraparteSeeking a counterpart
Acuerdos de soberanía de datos con institucionesData-sovereignty agreements with institutions
Por qué cuesta copiarloWhy it is hard to copy
Cada acuerdo se construye con confianza, uno por uno, y estimamos que lleva de 6 a 18 meses. No se puede comprar.Each agreement is built on trust, one at a time, and we estimate 6 to 18 months. It cannot be bought.
Qué existe hoyWhat exists today
Ninguno firmado. Buscamos una primera contraparte en El Salvador.None signed. We are looking for a first counterpart in El Salvador.
Qué lo erosionaríaWhat would erode it
Un gran proveedor que ofrezca nube dentro del país, un cambio político o un solo incidente.A large vendor offering in-country cloud, a political change, or a single incident.
04Construido, probado con datos sintéticosBuilt, tested on synthetic data
Verificación y evaluación propiasOur own verification and evaluation
Por qué cuesta copiarloWhy it is hard to copy
Un verificador marca lo que el modelo inventa y una prueba mide con intervalos de confianza. Por sí solo el código se puede rehacer; lo difícil es calibrarlo con habla regional y usarlo en resultados públicos.A verifier flags what the model invents and a scorer measures with confidence intervals. On its own the code can be rebuilt; what is hard to copy is calibrating it on regional speech and using it in public results.
Qué existe hoyWhat exists today
Un verificador con 182 pruebas automáticas, el medidor y el registro. Probados solo con datos sintéticos.A verifier with 182 automated tests, the scorer and the ledger. Tested only on synthetic data.
Qué lo erosionaríaWhat would erode it
Que otro equipo reescriba el código. Por eso no lo presentamos como foso por sí solo.Another team rewriting the code. That is why we do not present it as a moat on its own.
05Construido, sintéticoBuilt, synthetic
Mundos sintéticos revisados por expertosSynthetic worlds checked by experts
Por qué cuesta copiarloWhy it is hard to copy
El generador se puede copiar. Lo difícil de copiar son los atlas regionales revisados por personas expertas y la red de revisores.The generator can be copied. What is hard to copy are the regional atlases checked by experts and the network of reviewers.
Qué existe hoyWhat exists today
Un generador de reuniones ficticias de la región, con actas correctas conocidas. La revisión de expertos no ha empezado.A generator of fictional regional meetings with known correct minutes. Expert review has not started.
Qué lo erosionaríaWhat would erode it
Que modelos de frontera produzcan datos parecidos a bajo costo. Además, los datos sintéticos enseñan nuestras plantillas, no el habla real.Frontier models producing similar data cheaply. Also, synthetic data teaches our templates, not real speech.
El ciclo que queremos poner en marchaThe loop we want to set turning
Línea discontinua: todavía no existe. Es el diseño de un ciclo, no uno que ya gire. Cada vuelta solo cuenta si cada paso se puede comprobar por fuera.Dashed line: does not exist yet. It is the design of a loop, not one that already turns. Each turn only counts if each step can be checked from outside.
Qué NO es nuestro fosoWhat is NOT our moat
Los modelos abiertosOpen modelsCualquiera puede descargarlos, y el siguiente será gratis para todos. Los usamos; no son nuestros.Anyone can download them, and the next one will be free for everyone. We use them; they are not ours.
Los librosBooksLeer un libro en voz alta no enseña cómo habla la región, y la lectura es una copia del libro.Reading a book aloud does not teach how the region speaks, and the reading is a copy of the book.
El texto de la webWeb textLos modelos base ya lo traen, y buena parte tiene derechos de autor. Lo que cualquiera puede bajar, otro también.Base models already contain it, and much of it is under copyright. Whatever anyone can download, someone else can too.
Un foso empieza el día en que la primera grabación con consentimiento entra al registro y alguien que no entrena modelos sella el primer examen.A moat starts the day the first consented recording enters the ledger and someone who does not train models seals the first exam.
04 / Marco ético04 / Ethical framework
Ocho principios que rigen el programa.Eight principles that govern the program.
Estos principios aplican a los datos, al entrenamiento y al uso de Vocero. Cada uno se desarrollará en un documento completo. El borrador se comparte a solicitud.
These principles apply to the data, the training and the use of Vocero. Each will be set out in a full document. The draft is shared on request.
01
Dignidad y consentimiento informado
Dignity and informed consent
Nadie aporta su voz sin saber para qué se usa. El consentimiento es escrito, claro y específico.
No one contributes their voice without knowing what it is for. Consent is written, clear and specific.
BorradorDraftEl marco ético completo está disponible como borrador para revisión legal. No es asesoría legal y todavía no lo ha revisado asesoría salvadoreña.The full ethical framework is available as a draft for legal review. It is not legal advice and has not yet been reviewed by Salvadoran counsel.Leer el marco éticoRead the ethical framework
In preparationEl documento completo del marco ético está en preparación. El borrador se comparte a solicitud, escribiendo a hola@vocero.icu. Todavía no publicamos su texto.The full ethical framework document is in preparation. The draft is shared on request, by writing to hola@vocero.icu. We have not published its text yet.
05 / Estado05 / Status
Where we are today.
Vocero is in development and used only inside CareCompile. Dictation works today with an engine that is not our own model. Our own model is not trained yet; that is why we publish no accuracy figures, customers or prices: there are none.
Status by component
Site and demonstrationsRunning, with synthetic data.Live today
Speech recognition and dictation in meetings todayRunning on an engine that is not our own model.Live today
Dictation rules layerBuilt and tested on text; speech recognition is not trained.Built
Consent-based voice recorderBuilt; closed until consent is approved.Built
Control consoleBuilt.Built
Ethics frameworkDraft for legal review.In preparation
Vocero-Bench test benchDesigned.In preparation
Open-model baselineNot measured yet.Not yet
Regional dataNo hours collected.Not yet
Our own modelNot trained.Not yet
Evidence ladder
1Today: in development and governed from the startWe are here
2Baseline measured
3Improvement measured in El Salvador
4Improvement across several countries with no regressions
5Best open model for Central American SpanishOnly with a sealed set covering six countries.
6Deployment in the country with a counterpart
We do not claim a rung until we have its evidence.
No accuracy figures: not measured yet.
No customer list: there is none.
No prices or launch dates.
Today's dictation uses an engine that is not our own model.
Last updated: October 6, 2026, version 3.2.0. This board is updated as the evidence changes.
06 / Contacto06 / Contact
Name a counterpart and let us open the conversation.
If your institution can designate a counterpart, or would rather see the demo with synthetic data first, tell us. We talk about voices and texts under the ethics framework before anything is recorded.
Para su equipo técnicoFor your technical teamAnexo técnicoTechnical annexArquitectura, dictado, entrenamiento, datos, evaluación, equipos y preguntas frecuentes. Es diseño y plan, no resultados medidos. Pulse para abrir.Architecture, dictation, training, data, evaluation, equipment and FAQ. It is design and plan, not measured results. Press to open.
Architecture
Two-model cascadeEar (speech to text) and brain (language)
Target language
Central American SpanishMeasured country by country, with voseo
Brain, target
~8 billion parametersQuantized to fit in ~12 GB
Own equipment
CareCompile own equipment24 GB to train, 12 GB to serve
Status
Nothing trained yetNo accuracy figures: not measured yet
A1 / ArquitecturaA1 / Architecture
Un oído, un cerebro y reglas que se pueden auditar.An ear, a brain, and rules you can audit.
Vocero es una cascada, no un único modelo de audio de punta a punta. Cada pieza tiene su propia métrica y se puede medir, corregir o reemplazar por separado. Los comandos de dictado los aplica un motor de reglas, no un modelo de lenguaje. Las salidas pasan por una verificación contra la transcripción y, al final, por una persona.
Vocero is a cascade, not a single end-to-end audio model. Each part has its own metric and can be measured, fixed or replaced on its own. Dictation commands are applied by a rule engine, not by a language model. Outputs are checked against the transcript and, at the end, by a person.
Pase el cursor o use Tab sobre un cuadro: qué hace, cómo se mide y en qué estado está.Hover or Tab onto a box: what it does, how it is measured and its status.
Objetivo, no logradoTarget, not achieved
Basado en reglasRule-based
Se medirá despuésMeasured later
PersonaHuman
1Punto de medición contra el conjunto selladoMeasurement point against the sealed set
AudioAudioMicrófono para dictar, o el audio de una reunión.A microphone for dictation, or meeting audio.
Detección de vozVoice activity detectionSepara la voz del silencio antes de transcribir.Separates speech from silence before transcription.
Oído: voz a texto con puntuaciónEar: speech to text with punctuationModelo abierto de reconocimiento de voz, a ajustar con audio regional.An open speech recognition model, to be fine-tuned on regional audio.
Modo dictado: capa de comandosDictation mode: command layerReglas deterministas aplican coma, punto, nueva línea y correcciones. Resultado: texto dictado.Deterministic rules apply comma, period, new line and corrections. Result: dictated text.
Modo reunión: transcripciónMeeting mode: transcriptCon hablantes y marcas de tiempo, más un glosario regional y la lista de asistentes.With speakers and timestamps, plus a regional glossary and the attendee list.
Cerebro: modelo de lenguajeBrain: language modelObjetivo de unos 8 mil millones de parámetros, cuantizado para unos 12 GB. Escribe actas, tareas y respuestas con cita.Target of about 8 billion parameters, quantized for about 12 GB. Writes minutes, tasks and cited answers.
VerificaciónGrounding checkCada nombre, cifra, fecha y decisión debe aparecer en la transcripción; si no, se marca.Every name, number, date and decision must appear in the transcript; otherwise it is flagged.
Revisión humanaHuman reviewUna persona revisa antes de guardar o enviar.A person reviews before anything is saved or sent.
Conjunto de evaluación selladoSealed evaluation setMide cada pieza por separado. Lo custodia alguien que no entrena modelos.Measures each part on its own. Held by someone who does not train models.
Pieza 1Part 1Candidates
Oído
Ear
Reconocimiento de voz con puntuación automática en reuniones. En dictado, escribe los comandos como etiquetas (<coma>, <nueva_linea>) en lugar de adivinar.
Speech recognition with automatic punctuation in meetings. In dictation, it writes commands as tags (<coma>, <nueva_linea>) instead of guessing.
Pieza 2Part 2Rules
Capa de comandos
Command layer
Un motor de reglas probado: inserta puntuación, borra o corrige, y todo se puede deshacer. «Entró en coma» no es una coma: solo cuenta como comando en una pausa.
A tested rule engine: inserts punctuation, deletes or corrects, and every action can be undone. "Entró en coma" is not a comma: a word counts as a command only at a pause.
Pieza 3Part 3Candidates
Cerebro
Brain
Lee la transcripción y devuelve datos estructurados: acta, tareas con responsable y fecha, y respuestas con la marca de tiempo de donde salen.
Reads the transcript and returns structured data: minutes, tasks with owner and due date, and answers with the timestamp they come from.
Pieza 4Part 4Always
Verificación y persona
Check and a person
Lo que no coincide con la transcripción se marca, no se presenta como hecho. Vocero asiste; una persona revisa y decide.
Anything that does not match the transcript is flagged, not shown as fact. Vocero assists; a person reviews and decides.
A2 / Demostración: dictadoA2 / Demo: dictation
Speak, and it is written.
In dictation mode, the ear writes spoken commands as tags and a rule layer applies them. Hold the button and watch the raw text become final text.
Synthetic and simulated: does not use your microphone.
Dictation commands
Spoken
Effect
«coma»
types ,
«punto»
types . and capitalizes next
«nueva línea»
line break
«no perdón»
corrects the previous word
Done. See how the comma, period, new line and correction were applied.
What the ear transcribesestimadosmiembrosdelcomitécomabuenastardespuntonuevalíneaelinformeestarálistoellunesnoperdónmartespunto
Final text, after the command layer
Estimados miembros del comité, buenas tardes.
El informe estará listo el martes.
The sample is a synthetic script; the command layer that processes it is real and runs in your browser. A person should review the text before it is saved.
A3 / Cómo se entrenaA3 / How it is trained
Cinco etapas. Cada una con una puerta.Five stages. Each one has a gate.
No entrenamos desde cero: eso requiere miles de GPU. Partimos de modelos abiertos y los ajustamos con datos regionales. Ninguna etapa avanza si no pasa su puerta, y no hay fechas fijas: cada etapa espera a la anterior. Hoy ninguna etapa se ha ejecutado.
We do not train from scratch: that takes thousands of GPUs. We start from open models and adapt them with regional data. No stage moves on until it passes its gate, and there are no fixed dates: each stage waits for the one before. Today no stage has been run.
0
Línea base
Baseline
Correr los modelos base sin tocar, y el motor de dictado que funciona hoy, sobre el conjunto sellado. Error por país y puntaje de actas: la cifra que hay que superar.
Run the untouched base models, and the dictation engine running today, on the sealed set. Error per country and minutes scores: the number to beat.
PuertaGateLínea base escrita, con fecha y versiones exactas de cada modelo.Baseline written down, with the date and exact model versions.
Pending
1
Datos
Data
Reunir, revisar licencias, limpiar y separar por país. El conjunto de evaluación se aparta y se sella antes de cualquier entrenamiento.
Collect, check licenses, clean and split by country. The evaluation set is set aside and sealed before any training.
PuertaGateCada fuente con nota de licencia. Conjunto sellado con huella SHA-256.Every source has a license note. Set sealed with a SHA-256 fingerprint.
Pending
2
Oído
Ear
Ajuste fino de un modelo abierto de voz con audio regional, mezclado con español general para que no olvide lo que ya sabe.
Fine-tune an open speech model on regional audio, mixed with general Spanish so it does not forget what it already knows.
PuertaGateMejor que la línea base en el conjunto sellado, sin empeorar en ningún país, y al menos igual al dictado actual.Better than baseline on the sealed set, not worse in any country, and at least as good as today's dictation.
Pending
3
Cerebro
Brain
Ajuste por instrucciones (QLoRA) para actas, tareas y respuestas con cita; después, ajuste por preferencias contra invenciones. Preentrenamiento continuo solo si hace falta.
Instruction tuning (QLoRA) for minutes, tasks and cited answers; then preference tuning against invented content. Continued pretraining only if needed.
PuertaGateLos revisores lo califican mejor que el modelo base, sin nombres, cifras ni decisiones inventados.Reviewers rate it above the base model, with no invented names, numbers or decisions.
Pending
4
Paquete
Package
Fusionar los adaptadores, cuantizar y servir en equipos propios, detrás de un interruptor apagado por defecto.
Merge the adapters, quantize, and serve on our own equipment, behind a switch that is off by default.
PuertaGateAprobación explícita del responsable del programa para encender el interruptor.Explicit approval from the program lead to turn the switch on.
Pending
A4 / Preentrenamiento y datosA4 / Pretraining and data
El cuello de botella son los datos regionales.The bottleneck is regional data.
Los conjuntos abiertos de voz en español vienen sobre todo de España, México y Sudamérica. Hay muy poco audio abierto de Centroamérica, y sin él el oído no mejora donde importa. Por eso la mayor parte del trabajo real es reunir datos regionales con consentimiento.
Open Spanish speech sets come mostly from Spain, Mexico and South America. There is very little open Central American audio, and without it the ear will not improve where it matters. That is why most of the real work is collecting consented regional data.
Preentrenamiento continuoContinued pretrainingCPT
Seguir entrenando un modelo ya preentrenado con mucho texto sin etiquetar de la región: documentos públicos, noticias, enciclopedias con licencia. El modelo aprende vocabulario, nombres y giros. No le enseña una tarea. Lo haremos solo si la línea base muestra un vacío de vocabulario que el glosario no resuelve.
Keep training an already pretrained model on a large amount of unlabeled regional text: public documents, news, licensed encyclopedias. The model learns vocabulary, names and turns of phrase. It does not learn a task. We will do it only if the baseline shows a vocabulary gap that the glossary does not fix.
Ajuste finoFine-tuningLoRA / QLoRA
Entrenar con ejemplos etiquetados de la tarea: audio con su transcripción para el oído; transcripción con su acta revisada para el cerebro. Con LoRA solo se entrenan unas matrices pequeñas añadidas al modelo; por eso cabe en una GPU de 24 GB. Un ajuste completo de 8B necesitaría unos 128 GB.
Train on labeled examples of the task: audio with its transcript for the ear; a transcript with its reviewed minutes for the brain. With LoRA only small added matrices are trained, which is why it fits on a 24 GB GPU. A full fine-tune of an 8B model would need about 128 GB.
Qué hace distinto al español de El SalvadorWhat makes Salvadoran Spanish different
Voseo: «vos tenés», «decime». Un modelo entrenado con otro español puede «corregirlo» sin pedirlo.Voseo: "vos tenés", "decime". A model trained on other Spanish may "correct" it unasked.
Aspiración de la /s/ final: la /s/ al final de sílaba suena como [h], lo que confunde al reconocimiento de voz general.Final /s/ aspiration: syllable-final /s/ sounds like [h], which confuses general speech recognition.
Vocabulario local: «pisto», «cipote», «chunche», y nombres de lugares, instituciones y apellidos de la región.Local vocabulary: "pisto", "cipote", "chunche", and regional place names, institutions and family names.
Reuniones reales: cambios entre español e inglés, varios hablantes, audio de teléfono.Real meetings: switching between Spanish and English, several speakers, phone-quality audio.
Necesidades de datos para El SalvadorData needs for El SalvadorPlanning estimate
Data
What for
Amount
Possible source
Rule
Training audio with transcripts
Adapt the ear to Salvadoran speech
~20 hplanning estimate
Consenting speakers: scripted reading and spontaneous speech; every region, age and gender
Written consent stating the audio trains a model
Sealed test audio
Measure without cheating
~3 hplanning estimate
Speakers different from the training speakers
Never used for training; held by a custodian
Dictation clips
Commands, numbers, dates and regional names
to be setdepends on the script
Scripts we write, read by consenting speakers
Same consent as the training audio
Regional text
Vocabulary and names for the brain
to be setonly if the baseline needs it
Public documents and news archives whose license allows it
License note per source; no scraping of sites that forbid it
Meeting-style examples
Teach minutes, tasks and answers
thousandsfor the whole program
Synthetic meetings we write; real ones only with consent
Labelled synthetic; a person approves each example
General Spanish
Keep the ear from forgetting what it knows
existingopen sets
Open Spanish speech sets
License reviewed and recorded before use
Hours are planning estimates, not commitments or data already collected. No audio has been recorded for this program yet.
A5 / EvaluaciónA5 / Evaluation
Vocero-Bench CentroaméricaCentral America.
Un banco de pruebas diseñado para medir dictado, transcripción, actas y respuestas por país, calificado a ciegas y con intervalos de confianza. Está en diseño: no se ha grabado audio, no se ha sellado el conjunto y no se ha calificado ningún modelo.
A benchmark designed to measure dictation, transcription, minutes and answers per country, scored blind and with confidence intervals. It is in design: no audio has been recorded, no set has been sealed and no model has been scored.
Word error rate
WER = (S + D + I) / N
Substitutions, deletions and insertions over reference words, at corpus level. Published with and without accents.
Character error rate
CER
The same formula over characters. Useful with proper names and local words.
Punctuation
F1 = 2PR / (P + R)
Commas, periods, line breaks, and the opening marks ¿ and ¡ as classes of their own.
Commands
LCS(ref, hip) / Nref
Command accuracy in order, false activations, and commands typed out instead of executed.
Reviewer rubric
1 a 5 · κ ≥ 0.70
Two native reviewers per country, blind: completeness, correctness, usefulness and register. Disagreements are published.
Hallucinations
k de n, cota 95%
Any name, number or decision not in the transcript. We never say "zero".
Sealed set80% sealed test and 20% public development, split by speaker and by meeting.
Independent custodianA person who does not train models holds the test audio. Trainers never see it; they only get per-country aggregates.
Public fingerprintEvery file has a SHA-256; the manifest hash identifies the set and is published before training.
set-id: sha256(MANIFEST.tsv) = pending, the set does not exist yet
Contamination checksBefore every training run: hashes, speaker identity and 13-word overlaps against the test set.
Results by country: every cell empty because nothing has been measured
Country
WER
CER
Punct. F1
Commands
Minutes rubric
Hallucinations
SVEl Salvador
not measured
not measured
not measured
not measured
not measured
not measured
GTGuatemala
not measured
not measured
not measured
not measured
not measured
not measured
HNHonduras
not measured
not measured
not measured
not measured
not measured
not measured
NINicaragua
not measured
not measured
not measured
not measured
not measured
not measured
CRCosta Rica
not measured
not measured
not measured
not measured
not measured
not measured
PAPanama
not measured
not measured
not measured
not measured
not measured
not measured
Cells are empty on purpose: no model has been measured yet. When there are results, they will be published with the benchmark version, the date, the list of systems compared and the method.
A6 / Equipos propiosA6 / Our own equipment
Dimensionado para nuestras GPU.Sized for our own GPUs.
Entrenamos y servimos en equipos propios de CareCompile. Las cifras de memoria son objetivos y estimaciones aritméticas, no mediciones; la etapa 0 incluye una prueba corta para confirmarlas antes de cualquier trabajo largo.
We train and serve on CareCompile's own equipment. The memory figures are targets and arithmetic estimates, not measurements; stage 0 includes a short dry run to confirm them before any long job.
Training
Training machine · 24 GB
QLoRA tuning of a ~8B model and fine-tuning of the ear. Estimate for a typical run: about 12 to 15 GB of 24.
4-bit
012 GB24 GB
Base weights, 4-bit
~5.0 GB
LoRA adapters and optimizer
~1.0 GB
Activations, 8K tokens
~4-7 GB
CUDA context
~1.5 GB
Serving
Serving machine · 12 GB
Target: brain at 5 bits with an 8-bit cache, next to the ear, on one card. It fits, but only just: about 10 GB of 12.
5-bit
06 GB12 GB
Brain, 5-bit weights
~5.7 GB
8-bit KV cache, 16K tokens
~1.3 GB
CUDA context
~0.8 GB
Ear
~1.5-2.5 GB
Candidates to evaluate
Open models
Chosen from the baseline, per country, on our sealed set. No candidate has any results from us yet.
Ear: open speech models
Whisper large-v3-turbo OpenAI · MIT
Parakeet TDT 0.6B v3 NVIDIA · CC-BY-4.0
Granite Speech 4.1 2B IBM · Apache 2.0
Brain: open-weight base models
Granite 4.2 8B IBM · Apache 2.0
Gemma 4 E4B Google · Apache 2.0
Other open models of similar size
A7 / PreguntasA7 / Questions
Frequently asked questions.
What is Vocero?
Vocero is CareCompile's own AI. It dictates, listens to meetings, summarizes them into minutes and tasks, and answers questions about what was said, with the line each answer comes from.
Is your own model trained yet?
No. There is a staged training plan and an evaluation design, but no stage has been run. The dictation that works today uses another engine. We publish no accuracy figures because there are none.
Which models will you use?
Open models that are still candidates to evaluate: open speech models for the ear and open-weight base models of about 8 billion parameters for the brain. They will be chosen from baseline results on our sealed set, country by country.
Why not use a general cloud model?
General models are trained mostly on other Spanish and are not measured country by country in Central America. We also want audio to stay on our own equipment. A general model will still reason better at many tasks; our goal is narrow and measurable: regional Spanish, dictation and minutes.
Where does it run?
On CareCompile's own equipment: one training machine with 24 GB of GPU memory and one serving machine with 12 GB, per the plan.
What happens to contributed data?
It is used only to build and evaluate Vocero, with written consent, on CareCompile equipment and never resold. Contributors can withdraw and ask for deletion by writing to hola@vocero.icu. We do not accept real patient data or private conversations without consent.
Which countries?
Vocero aims at the Spanish of Central America, and the benchmark is designed to measure each country separately. Coverage will be confirmed by measurement; we do not yet promise any particular country.
What data does this site keep?
This site uses no cookies and no analytics. The only thing your browser keeps is the language you choose. If you send the form, we keep what you write so we can reply. The sample dictation does not use your microphone. The details are in the privacy policy; for privacy and rights requests: hola@vocero.icu.
How do I request a demo?
Fill in the form at the end of this page. The demo uses synthetic data, not real meetings.
What is presentation mode?
Press the P key (or the P button in the top bar) for a simplified view: only the cover, the meeting demo, data for El Salvador, why this will be hard to copy, status and contact, in larger type. Made for projecting in a room. Press P again to exit.