
Telehealth was supposed to widen access. For patients who do not speak English, it narrowed access.
An analysis of California Health Interview Survey data published in Health Affairs found that patients with limited English proficiency used telehealth at a rate of 4.8 percent, against 12.3 percent among proficient English speakers, roughly half the odds even after adjusting for other sociodemographic factors.
A systematic review by researchers at the University of North Carolina’s Sheps Center for Health Services Research reached a starker finding: every one of the seven studies examining telehealth by modality reported significantly lower video use among patients with limited English proficiency. Of the roughly 25.6 million people in the United States with limited English proficiency, those who do connect are disproportionately routed to audio-only care, the modality where interpretation is hardest and clinical observation is weakest.
Here is what competence looks like in practice.
Book the language, not just the appointment
Language needs must be captured as a structured field at scheduling and must carry an interpreter reservation with it. Two failures are near-universal. First, preferred language is recorded but never triggers anything downstream. Second, the language is recorded too coarsely “Chinese” instead of Cantonese or Mandarin, “Arabic” without dialect, “Spanish” without noting an Indigenous first language behind it. A mismatched interpreter burns the appointment slot as thoroughly as no interpreter at all.
Book by encounter type as well as language. A 15-minute medication check and a 45-minute oncology consult are different assignments, and rare languages need lead time that on-demand queues cannot always deliver.
Send three links, and test them
The interpreter needs a dedicated participant link not the patient’s link forwarded, and not a phone line bridged into a video visit. Bridging is where three-way encounters most often collapse: the interpreter cannot see the patient, and roughly half of what an interpreter relies on disappears.
That is not only a quality problem. The technical standards carried into Section 1557 require a sharply delineated image large enough to show the faces of both the interpreter and the person being served, along with real-time video free of lags or irregular pauses. The rule explicitly bars reliance on low-quality video remote interpreting, and OCR’s December 2024 Dear Colleague letter signaled that these obligations are active and enforceable. An interpreter on a phone line is a compliance exposure, not a workaround.
Where the platform allows it, pin the interpreter’s video so the patient can see them throughout, and confirm the connection before the patient joins.
Spend sixty seconds on the pre-session
The single highest-yield habit in interpreted care costs one minute. Before the clinical conversation begins, the clinician and interpreter align on the encounter: the visit’s purpose, any specialized terminology likely to come up, whether difficult news is expected, and how the interpreter should signal a need to pause.
This is also where the interpreter states the ground rules to the patient that everything said will be interpreted, that nothing will be omitted or added, and that the encounter is confidential. Patients who understand this speak more freely. Patients who do not tend to compress their answers.
Run the visit in the first person
Speak directly to the patient, not to the interpreter. “Where does it hurt?” not “Ask her where it hurts.” The interpreter renders in the first person. This is standard practice, and it is not a stylistic nicety: third-person framing pushes the patient out of their own encounter and invites the clinician to build rapport with the wrong person.
Deliver one or two sentences at a time and pause. Long uninterrupted passages force the interpreter to summarize, and summarization is where clinical detail is lost. Expect the interpretation to take longer than the original. Build the extra time into the slot rather than accelerating to recover it.
Two habits matter especially on video. Silence carries no visual explanation, so narrate what you are doing when you turn to the chart. And when the interpreter requests a clarification, treating it as a quality signal rather than an interruption it usually means a term has no clean equivalent, which is precisely the moment a misunderstanding would otherwise pass unnoticed.
Document the interpretation, not just the visit
Record the interpreter’s name or ID, the language and dialect, the modality, and the duration. Most organizations cannot currently stratify outcomes by language because this data never reaches the chart, which makes a known disparity invisible to the quality committee that is accountable for it.
Teach-back is where a language-discordant visit is verified or lost. Ask the patient to explain the plan in their own words, through the interpreter, before anyone disconnects. Written follow-up should be reviewed by a qualified human, the 2024 rule requires exactly that for critical communications produced with machine translation.
The system-level point
None of this is difficult. It is unassigned. Interpreter integration usually belongs to no one: not the telehealth platform team, not the clinical informatics group, not the compliance office that owns the language access plan on paper.
Health systems now buying ambient documentation, AI translation, and virtual-first care are making architectural decisions about a population that is already using video least. Deciding who owns the three-way visit before those tools are configured rather than after is the cheapest equity investment on the roadmap.
About Philip Rosen
About Philip Rosen
Philip Rosen is the Founder and Managing Director of Capital Linguists, a US-based language services firm providing certified interpreting and translation across more than 210 languages. He works with healthcare systems, government agencies, and academic institutions on language access in high-stakes settings.
