Omnimed AI

    Omni-AI Integration Guide

    Integrate Omni-AI medical transcription into your product in minutes.

    1. Widget

    Drop a <script> tag into your page. A floating recording button is injected; the generated report can be auto-inserted into the active editor.

    Use this if: you want the fastest integration and are happy with the default recording UI.

    2. Direct API

    POST an audio file to /api/omni-ai and handle the response yourself.

    Use this if: you need full control over the recording UI or are integrating from a native/desktop client.

    Base URL: https://api.omniai.com.ar

    Overview

    Omni-AI turns a medical dictation (audio) into a structured, template-filled report in a single request. You can integrate in two ways:

    1. Embeddable Widget — drop a <script> tag into your page. The doctor records from a floating button, the report is returned and can be auto-inserted into your editor.
    2. Direct API — call POST https://api.omniai.com.ar/api/omni-ai with the audio file yourself and handle the response.

    Both routes use the same omni-ai endpoint under the hood. Start with the Widget unless you need full control of the recording UI.

    Authentication

    Each integration partner receives a medical_center_id (provisioned by Omni). The center ID identifies the templates, language and report style used to generate the output. No bearer token is required today.

    Contact [email protected] to provision a center ID.

    Widget

    Embed

    <script src="https://api.omniai.com.ar/static/widget/v1/omni-widget.min.js"
            data-center-id="YOUR_CENTER_ID"
            async></script>

    A floating recording button is injected in the bottom-right. The doctor clicks it, dictates, stops, and the widget shows the generated report.

    Auto-insert the report

    The widget emits an omniai:report event when the doctor clicks "Insertar". Listen for it and paste the HTML (or plain text) into your editor:

    window.addEventListener('omniai:report', (event) => {
      const { html, text } = event.detail;
      // e.g. paste into the focused contenteditable / rich-text field
      document.execCommand('insertHTML', false, html);
    });

    Or initialize manually with a callback:

    OmniWidget.init({
      centerId: 'YOUR_CENTER_ID',
      onReport: ({ html, text }) => {
        // Insert into your report editor
      }
    });

    Widget lifecycle

    StateMeaning
    idleReady to record
    recordingMicrophone active
    pausedRecording paused
    processingAudio uploaded, report generating
    resultReport shown in panel

    Browser requirements

    • HTTPS (required for microphone access)
    • Permission granted for navigator.mediaDevices.getUserMedia
    • Max recording duration: 10 minutes (configurable)

    Direct API

    Generate a report

    POST https://api.omniai.com.ar/api/omni-ai
    Content-Type: multipart/form-data

    Parameters

    ParameterTypeRequiredDescription
    fileFileConditionalAudio file (wav, mp3, m4a, ogg, flac, webm). Required unless transcription_text is sent
    transcription_textStringConditionalDictation text used instead of audio (see below). Required unless file is sent
    medical_center_idStringYesCenter identifier provided by Omni
    user_idStringNoDoctor/user identifier for analytics
    languageStringNoLanguage code (default spa)
    metadataString (JSON)NoFree-form JSON metadata, max 10KB
    dicom_metadataString (JSON)NoDICOM metadata for study fallback
    template_htmlString (HTML)NoPreselected template: generate the report against exactly this template (see below)
    template_nameStringNoLabel for the preselected template; echoed back as study_name / template_used
    include_templateBooleanNoReturn the template HTML the server used, for diff highlighting
    include_rtfBooleanNoReturn RTF version of each report

    Note: Without template_html, study detection is automatic from the transcription and the template is matched against the center's stored templates — you do not need to send the study name.

    Sending text instead of audio (transcription_text)

    If your system already has the dictation as text (typed, or transcribed on your side), send it as transcription_text and omit the audio — the transcription step is skipped and the text enters the pipeline exactly where the transcription would: study detection, template matching and generation behave identically. It combines naturally with template_html (no transcription and no matching — the fastest, fully deterministic path).

    • Audio always wins. If you send both, the audio is transcribed and transcription_text is ignored entirely. There is no fallback: if the audio fails to transcribe (e.g. muted microphone) the call errors even when text was sent.
    • A blank file field or a blank transcription_text ("" or whitespace) is treated as absent. Sending neither returns a 400.
    • Plain text, max 50KB.
    • In the response: duration_seconds is 0.0, speaker_count is 1, language_detected echoes your language parameter, and metadata.input_source is "text" (vs "audio").
    • Billing is unchanged — a report generated from text consumes credit exactly like one generated from audio.

    Sending your own template (template_html)

    When your system knows which template the report must follow (per clinic, per referring center, per modality), send it in the call:

    • Study detection and template matching are skipped entirely — the report deterministically follows your template.
    • The result is always a single report, even if the dictation mentions several studies.
    • The HTML is sanitised server-side before use; max size is 100KB. Oversized HTML, or markup that is empty after sanitisation, returns a 400. A blank field ("" or whitespace) is treated as absent: the call falls back to the normal matching path.
    • HTML is the recommended format: it preserves bold, headings and block structure.
    • In the response: template_match_confidence is 1.0 and metadata.detection_source is "preselected_template".
    • With include_template=true, template_html in the response is the sanitised template the server actually used — diff against this, not your original.

    Success response

    {
      "success": true,
      "message": "Successfully processed audio and generated 1 report(s)",
      "data": {
        "transcription": "...",
        "language_detected": "spa",
        "duration_seconds": 42.1,
        "speaker_count": 1,
        "reports": [
          {
            "study_name": "IRM de Rodilla",
            "template_used": "IRM de Rodilla",
            "template_match_confidence": 1.0,
            "report_html": "<p>...</p>",
            "report_rtf": null,
            "transcription_segment": "...",
            "template_html": null
          }
        ],
        "processing_time_ms": 7441,
        "metadata": {
          "study_id": 12345,
          "detection_source": "preselected_template",
          "...": "..."
        }
      }
    }

    Tip: Keep data.metadata.study_id if you plan to refine this report later.

    Examples

    cURL

    curl -X POST https://api.omniai.com.ar/api/omni-ai \
      -F "[email protected]" \
      -F "medical_center_id=YOUR_CENTER_ID" \
      -F "user_id=dr_martinez" \
      -F "language=spa"

    cURL (preselected template)

    curl -X POST https://api.omniai.com.ar/api/omni-ai \
      -F "[email protected]" \
      -F "medical_center_id=YOUR_CENTER_ID" \
      -F "template_html=<h2>IRM de Rodilla</h2><p>...</p>" \
      -F "template_name=IRM de Rodilla" \
      -F "include_template=true"

    Python

    import requests
    
    url = "https://api.omniai.com.ar/api/omni-ai"
    
    with open("audio.m4a", "rb") as f:
        response = requests.post(
            url,
            files={"file": f},
            data={
                "medical_center_id": "YOUR_CENTER_ID",
                "user_id": "dr_lopez",
                "language": "spa",
            },
        )
    
    response.raise_for_status()
    result = response.json()
    
    for report in result["data"]["reports"]:
        print(report["study_name"], "->", report["report_html"])

    JavaScript (fetch)

    async function processAudio(audioFile, centerId, userId) {
      const form = new FormData();
      form.append('file', audioFile);
      form.append('medical_center_id', centerId);
      form.append('user_id', userId);
      form.append('language', 'spa');
    
      const res = await fetch('https://api.omniai.com.ar/api/omni-ai', {
        method: 'POST',
        body: form,
      });
    
      if (!res.ok) {
        const err = await res.json().catch(() => ({}));
        throw new Error(err.error || `HTTP ${res.status}`);
      }
    
      const { data } = await res.json();
      return data.reports[0].report_html;
    }

    Refine a report

    POST https://api.omniai.com.ar/api/omni-ai/refine
    Content-Type: multipart/form-data

    The doctor dictates a correction to an already-generated report ("cambiar menisco interno por menisco externo"). The server transcribes the audio, changes only the targeted element — plus the report's Conclusion when the change affects it — and returns the updated report. Refinement never regenerates the whole report, and it is not billed.

    Parameters

    ParameterTypeRequiredDescription
    fileFileConditionalAudio with the dictated correction. Required unless transcription_text is sent
    transcription_textStringConditionalThe correction as text instead of audio (max 50KB, same rules as on generation: audio wins, blank = absent)
    medical_center_idStringYesCenter identifier provided by Omni
    report_htmlString (HTML)YesThe last generated report, exactly as returned by the API (or as last refined)
    template_htmlString (HTML)NoThe original template of the report; used as a read-only style/format reference (max 100KB)
    template_nameStringNoTemplate/study-type name, enables per-study-type rules
    selected_htmlString (HTML)NoExact HTML element to refine; omit to let the model locate the target from the dictation
    study_idIntegerNometadata.study_id from the generation call, to track the refinement
    user_idStringNoDoctor/user identifier for analytics
    languageStringNoLanguage code (default spa)

    Success response

    {
      "success": true,
      "message": "Successfully refined report",
      "data": {
        "transcription": "cambiar menisco interno por menisco externo",
        "report_html": "<p>...full report with the correction applied...</p>",
        "patch": {
          "refined_text": "<p>Menisco externo sin alteraciones.</p>",
          "target_html": "<p>Menisco interno sin alteraciones.</p>",
          "conclusion_text": null,
          "conclusion_original_text": null
        },
        "processing_time_ms": 3120,
        "metadata": {
          "medical_center_id": "your-center",
          "study_id": 12345,
          "event_id": "678",
          "model_used": "gemini-2.5-flash",
          "...": "..."
        }
      }
    }
    • data.report_html is the full report with the correction already applied — use it directly.
    • data.patch describes the exact regions that changed, if you prefer to splice them into your own editor: replace the first occurrence of target_html with refined_text, and of conclusion_original_text with conclusion_text (when non-null).
    • To chain refinements, send the returned report_html as the input of the next refine call.

    Example

    curl -X POST https://api.omniai.com.ar/api/omni-ai/refine \
      -F "[email protected]" \
      -F "medical_center_id=YOUR_CENTER_ID" \
      -F "report_html=<p>...last generated report...</p>" \
      -F "template_html=<p>...original template...</p>" \
      -F "study_id=12345"

    Error response

    Errors return a non-2xx status with:

    {
      "success": false,
      "error": "Human-readable error message",
      "error_code": 400
    }
    StatusCauseFix
    400Missing input (neither file nor transcription_text), unsupported audio format, malformed metadata JSON, or oversized template_htmlSend wav/mp3/m4a/ogg/flac/webm or transcription_text (max 50KB); validate JSON under 10KB; keep template_html under 100KB
    403Center disabled / not provisionedContact Omni to enable your medical_center_id
    500Transcription or report generation failedRetry with backoff; contact support if persistent

    Highlighting AI changes (client-side diff)

    The Widget highlights what the AI changed by diffing the template against the generated report. You can do the same in your UI with jsdiff (diff on npm) — no server support needed:

    import { diffWords } from 'diff';
    
    function highlightChanges(templateText, reportText) {
      return diffWords(templateText, reportText)
        .map((part) =>
          part.added
            ? `<mark class="ai-change">${part.value}</mark>`
            : part.removed
              ? ''
              : part.value
        )
        .join('');
    }
    • Baseline: request include_template=true on generation and diff against the returned template_html.
    • After a refine, diff the previous report_html against the new one to flash only the correction.
    • Diff on visible text (e.g. from DOMParser) rather than raw HTML strings for cleaner highlights.

    Health check

    Use this endpoint to verify connectivity from your backend:

    curl https://api.omniai.com.ar/api/health

    Response:

    {
      "status": "healthy",
      "services": {
        "api": {
          "service": "api",
          "status": "healthy",
          "available": true
        }
      }
    }

    Support

    Questions and provisioning: [email protected]