Ear
A candidate open speech model would turn audio into text with speakers and times and produce tags for dictation commands.
Vocero separates speech, dictation rules, and meeting analysis so each component can be tested. The architecture, training, and benchmark form a technical plan; the site identifies separately what has already been built.
In development for internal use. Our own model is not trained yet.
The design avoids assigning to the language model work that belongs to speech recognition or explicit rules.
A candidate open speech model would turn audio into text with speakers and times and produce tags for dictation commands.
A non-AI layer applies punctuation, formatting, and corrections. This layer is built, and its demo version runs in the browser.
A candidate open base model would turn the transcript into decisions, tasks, entities, and answers with references.
The final output requires human review before it is saved or sent.
Every proposed stage has an evidence gate. No training stage for Vocero's own model has been run.
The plan compares open candidates on a sealed per-country set before selecting models or starting tuning.
The ear would be tuned on regional audio mixed with general Spanish. The brain would be tuned on minutes, tasks, and cited answers.
Regional text would be used for continued pretraining only if the baseline reveals a vocabulary gap that a glossary cannot solve.
The objective is to quantize and serve on CareCompile equipment, behind a switch that is off by default and requires explicit approval to activate.
The benchmark is designed to evaluate each country and separate development from testing. No audio has been recorded, no set sealed, and no model scored.
Planned metrics include word error rate, character error rate, punctuation, commands in order, and false activations.
Two native reviewers per country would score completeness, accuracy, usefulness, and register, and disagreements would be published.
Any name, number, or decision absent from the transcript counts as invented; a report would not claim zero without the corresponding statistical evidence.
The proposed method aims to prevent training from seeing or reproducing the test set.
The design assigns 80% to sealed testing and 20% to public development, with no speakers or meetings shared between them.
The person holding test audio does not train models; the training team would receive only aggregate results by country.
The plan publishes the SHA-256 manifest fingerprint and checks hashes, speaker identity, and thirteen-word overlaps before each run.
When results exist, they should include benchmark version, date, compared systems, and method. There are no figures to publish today.
Explore Vocero's design for transcribing meetings and producing minutes with decisions, tasks, and source citations. In development; the demo is synthetic.
Explore the scopeVocero investigates dictation for Salvadoran and Central American Spanish. Try the implemented command rules and explore the regional speech plan.
Explore the scopeCareCompile seeks an institutional counterpart in El Salvador to explore consented speech and text data, limited purpose, and independent evaluation.
Explore the scopeWe are looking for an institution willing to discuss regional meetings, voices, and texts. We explain purpose, consent, and governance before recording or receiving data.