Omni-AI Integration Guide
Integrate Omni-AI medical transcription into your product in minutes.
1. Widget
Drop a <script> tag into your page. A floating recording button is injected; the generated report can be auto-inserted into the active editor.
Use this if: you want the fastest integration and are happy with the default recording UI.
2. Direct API
POST an audio file to /api/omni-ai and handle the response yourself.
Use this if: you need full control over the recording UI or are integrating from a native/desktop client.
Base URL: https://api.omniai.com.ar
Overview
Omni-AI turns a medical dictation (audio) into a structured, template-filled report in a single request. You can integrate in two ways:
- Embeddable Widget — drop a
<script>tag into your page. The doctor records from a floating button, the report is returned and can be auto-inserted into your editor. - Direct API — call
POST https://api.omniai.com.ar/api/omni-aiwith the audio file yourself and handle the response.
Both routes use the same omni-ai endpoint under the hood. Start with the Widget unless you need full control of the recording UI.
Authentication
Each integration partner receives a medical_center_id (provisioned by Omni). The center ID identifies the templates, language and report style used to generate the output. No bearer token is required today.
Contact [email protected] to provision a center ID.
Widget
Embed
<script src="https://api.omniai.com.ar/static/widget/v1/omni-widget.min.js"
data-center-id="YOUR_CENTER_ID"
async></script>A floating recording button is injected in the bottom-right. The doctor clicks it, dictates, stops, and the widget shows the generated report.
Auto-insert the report
The widget emits an omniai:report event when the doctor clicks "Insertar". Listen for it and paste the HTML (or plain text) into your editor:
window.addEventListener('omniai:report', (event) => {
const { html, text } = event.detail;
// e.g. paste into the focused contenteditable / rich-text field
document.execCommand('insertHTML', false, html);
});Or initialize manually with a callback:
OmniWidget.init({
centerId: 'YOUR_CENTER_ID',
onReport: ({ html, text }) => {
// Insert into your report editor
}
});Widget lifecycle
| State | Meaning |
|---|---|
| idle | Ready to record |
| recording | Microphone active |
| paused | Recording paused |
| processing | Audio uploaded, report generating |
| result | Report shown in panel |
Browser requirements
- HTTPS (required for microphone access)
- Permission granted for
navigator.mediaDevices.getUserMedia - Max recording duration: 10 minutes (configurable)
Direct API
Generate a report
POST https://api.omniai.com.ar/api/omni-ai
Content-Type: multipart/form-dataParameters
| Parameter | Type | Required | Description |
|---|---|---|---|
| file | File | Conditional | Audio file (wav, mp3, m4a, ogg, flac, webm). Required unless transcription_text is sent |
| transcription_text | String | Conditional | Dictation text used instead of audio (see below). Required unless file is sent |
| medical_center_id | String | Yes | Center identifier provided by Omni |
| user_id | String | No | Doctor/user identifier for analytics |
| language | String | No | Language code (default spa) |
| metadata | String (JSON) | No | Free-form JSON metadata, max 10KB |
| dicom_metadata | String (JSON) | No | DICOM metadata for study fallback |
| template_html | String (HTML) | No | Preselected template: generate the report against exactly this template (see below) |
| template_name | String | No | Label for the preselected template; echoed back as study_name / template_used |
| include_template | Boolean | No | Return the template HTML the server used, for diff highlighting |
| include_rtf | Boolean | No | Return RTF version of each report |
Note: Without template_html, study detection is automatic from the transcription and the template is matched against the center's stored templates — you do not need to send the study name.
Sending text instead of audio (transcription_text)
If your system already has the dictation as text (typed, or transcribed on your side), send it as transcription_text and omit the audio — the transcription step is skipped and the text enters the pipeline exactly where the transcription would: study detection, template matching and generation behave identically. It combines naturally with template_html (no transcription and no matching — the fastest, fully deterministic path).
- Audio always wins. If you send both, the audio is transcribed and
transcription_textis ignored entirely. There is no fallback: if the audio fails to transcribe (e.g. muted microphone) the call errors even when text was sent. - A blank
filefield or a blanktranscription_text("" or whitespace) is treated as absent. Sending neither returns a 400. - Plain text, max 50KB.
- In the response:
duration_secondsis0.0,speaker_countis1,language_detectedechoes yourlanguageparameter, andmetadata.input_sourceis"text"(vs"audio"). - Billing is unchanged — a report generated from text consumes credit exactly like one generated from audio.
Sending your own template (template_html)
When your system knows which template the report must follow (per clinic, per referring center, per modality), send it in the call:
- Study detection and template matching are skipped entirely — the report deterministically follows your template.
- The result is always a single report, even if the dictation mentions several studies.
- The HTML is sanitised server-side before use; max size is 100KB. Oversized HTML, or markup that is empty after sanitisation, returns a 400. A blank field ("" or whitespace) is treated as absent: the call falls back to the normal matching path.
- HTML is the recommended format: it preserves bold, headings and block structure.
- In the response:
template_match_confidenceis1.0andmetadata.detection_sourceis"preselected_template". - With
include_template=true,template_htmlin the response is the sanitised template the server actually used — diff against this, not your original.
Success response
{
"success": true,
"message": "Successfully processed audio and generated 1 report(s)",
"data": {
"transcription": "...",
"language_detected": "spa",
"duration_seconds": 42.1,
"speaker_count": 1,
"reports": [
{
"study_name": "IRM de Rodilla",
"template_used": "IRM de Rodilla",
"template_match_confidence": 1.0,
"report_html": "<p>...</p>",
"report_rtf": null,
"transcription_segment": "...",
"template_html": null
}
],
"processing_time_ms": 7441,
"metadata": {
"study_id": 12345,
"detection_source": "preselected_template",
"...": "..."
}
}
}Tip: Keep data.metadata.study_id if you plan to refine this report later.
Examples
cURL
curl -X POST https://api.omniai.com.ar/api/omni-ai \
-F "[email protected]" \
-F "medical_center_id=YOUR_CENTER_ID" \
-F "user_id=dr_martinez" \
-F "language=spa"cURL (preselected template)
curl -X POST https://api.omniai.com.ar/api/omni-ai \
-F "[email protected]" \
-F "medical_center_id=YOUR_CENTER_ID" \
-F "template_html=<h2>IRM de Rodilla</h2><p>...</p>" \
-F "template_name=IRM de Rodilla" \
-F "include_template=true"Python
import requests
url = "https://api.omniai.com.ar/api/omni-ai"
with open("audio.m4a", "rb") as f:
response = requests.post(
url,
files={"file": f},
data={
"medical_center_id": "YOUR_CENTER_ID",
"user_id": "dr_lopez",
"language": "spa",
},
)
response.raise_for_status()
result = response.json()
for report in result["data"]["reports"]:
print(report["study_name"], "->", report["report_html"])JavaScript (fetch)
async function processAudio(audioFile, centerId, userId) {
const form = new FormData();
form.append('file', audioFile);
form.append('medical_center_id', centerId);
form.append('user_id', userId);
form.append('language', 'spa');
const res = await fetch('https://api.omniai.com.ar/api/omni-ai', {
method: 'POST',
body: form,
});
if (!res.ok) {
const err = await res.json().catch(() => ({}));
throw new Error(err.error || `HTTP ${res.status}`);
}
const { data } = await res.json();
return data.reports[0].report_html;
}Refine a report
POST https://api.omniai.com.ar/api/omni-ai/refine
Content-Type: multipart/form-dataThe doctor dictates a correction to an already-generated report ("cambiar menisco interno por menisco externo"). The server transcribes the audio, changes only the targeted element — plus the report's Conclusion when the change affects it — and returns the updated report. Refinement never regenerates the whole report, and it is not billed.
Parameters
| Parameter | Type | Required | Description |
|---|---|---|---|
| file | File | Conditional | Audio with the dictated correction. Required unless transcription_text is sent |
| transcription_text | String | Conditional | The correction as text instead of audio (max 50KB, same rules as on generation: audio wins, blank = absent) |
| medical_center_id | String | Yes | Center identifier provided by Omni |
| report_html | String (HTML) | Yes | The last generated report, exactly as returned by the API (or as last refined) |
| template_html | String (HTML) | No | The original template of the report; used as a read-only style/format reference (max 100KB) |
| template_name | String | No | Template/study-type name, enables per-study-type rules |
| selected_html | String (HTML) | No | Exact HTML element to refine; omit to let the model locate the target from the dictation |
| study_id | Integer | No | metadata.study_id from the generation call, to track the refinement |
| user_id | String | No | Doctor/user identifier for analytics |
| language | String | No | Language code (default spa) |
Success response
{
"success": true,
"message": "Successfully refined report",
"data": {
"transcription": "cambiar menisco interno por menisco externo",
"report_html": "<p>...full report with the correction applied...</p>",
"patch": {
"refined_text": "<p>Menisco externo sin alteraciones.</p>",
"target_html": "<p>Menisco interno sin alteraciones.</p>",
"conclusion_text": null,
"conclusion_original_text": null
},
"processing_time_ms": 3120,
"metadata": {
"medical_center_id": "your-center",
"study_id": 12345,
"event_id": "678",
"model_used": "gemini-2.5-flash",
"...": "..."
}
}
}data.report_htmlis the full report with the correction already applied — use it directly.data.patchdescribes the exact regions that changed, if you prefer to splice them into your own editor: replace the first occurrence oftarget_htmlwithrefined_text, and ofconclusion_original_textwithconclusion_text(when non-null).- To chain refinements, send the returned
report_htmlas the input of the next refine call.
Example
curl -X POST https://api.omniai.com.ar/api/omni-ai/refine \
-F "[email protected]" \
-F "medical_center_id=YOUR_CENTER_ID" \
-F "report_html=<p>...last generated report...</p>" \
-F "template_html=<p>...original template...</p>" \
-F "study_id=12345"Error response
Errors return a non-2xx status with:
{
"success": false,
"error": "Human-readable error message",
"error_code": 400
}| Status | Cause | Fix |
|---|---|---|
| 400 | Missing input (neither file nor transcription_text), unsupported audio format, malformed metadata JSON, or oversized template_html | Send wav/mp3/m4a/ogg/flac/webm or transcription_text (max 50KB); validate JSON under 10KB; keep template_html under 100KB |
| 403 | Center disabled / not provisioned | Contact Omni to enable your medical_center_id |
| 500 | Transcription or report generation failed | Retry with backoff; contact support if persistent |
Highlighting AI changes (client-side diff)
The Widget highlights what the AI changed by diffing the template against the generated report. You can do the same in your UI with jsdiff (diff on npm) — no server support needed:
import { diffWords } from 'diff';
function highlightChanges(templateText, reportText) {
return diffWords(templateText, reportText)
.map((part) =>
part.added
? `<mark class="ai-change">${part.value}</mark>`
: part.removed
? ''
: part.value
)
.join('');
}- Baseline: request
include_template=trueon generation and diff against the returnedtemplate_html. - After a refine, diff the previous
report_htmlagainst the new one to flash only the correction. - Diff on visible text (e.g. from
DOMParser) rather than raw HTML strings for cleaner highlights.
Health check
Use this endpoint to verify connectivity from your backend:
curl https://api.omniai.com.ar/api/healthResponse:
{
"status": "healthy",
"services": {
"api": {
"service": "api",
"status": "healthy",
"available": true
}
}
}Support
Questions and provisioning: [email protected]
