How Technology Is Changing Speech and Language Therapy: AI, Voice Analysis, AAC and Telepractice

Speech and language therapy is entering a new technological phase. Artificial intelligence, automatic speech recognition, voice analysis, augmentative and alternative communication, virtual reality, digital therapeutic activities and telepractice can extend how speech-language pathologists assess, treat and support communication.

But communication is more than producing the correct sound or achieving a higher score in an application.

The objective remains helping people communicate, understand, participate, express themselves and connect with others in real-life situations.

What is speech and language therapy?

Speech and language therapy—also known as speech-language pathology—is a health profession concerned with communication and swallowing across the lifespan.

Depending on the country and scope of practice, speech-language pathologists or speech and language therapists may assess and support people experiencing difficulties involving speech sound production, articulation, motor speech, language comprehension, language expression, fluency, voice, social communication, cognitive communication, literacy-related communication difficulties, augmentative and alternative communication, feeding and swallowing.

SLPs work with infants, children, adults and older people in hospitals, rehabilitation centers, schools, private practice, community services and increasingly through remote and hybrid models.

The precise scope of practice varies between professional and regulatory systems.

From face-to-face interaction to digitally supported communication

Speech and language therapy has always depended heavily on human interaction.

Traditionally, clinicians have relied on direct observation, conversation, standardized assessment, repetition, modeling, visual materials, toys, pictures, written exercises and home practice.

Technology gradually expanded this toolkit through audio recording, video, computerized assessment, communication devices, speech-generating devices and specialized applications.

Today, the next stage includes artificial intelligence, automatic speech recognition, machine learning, advanced AAC, real-time voice analysis, virtual and mixed reality, multi-user virtual environments, telepractice, mobile applications, digital games and remote patient monitoring.

Why can technology be particularly relevant to speech and language therapy?

Many communication interventions require repetition, frequent practice, immediate feedback, individualized stimuli, progression in difficulty, practice outside the therapy session, family or caregiver involvement, opportunities for spontaneous interaction and practice in realistic social situations.

These characteristics make speech and language therapy particularly compatible with well-designed digital tools.

Technology can potentially increase the number of practice opportunities between appointments and create communication situations that are difficult to reproduce inside a therapy room.

But increased practice is useful only if the target is appropriate, the feedback is sufficiently accurate, the person understands the activity, the interaction remains meaningful and the task transfers to everyday communication.

A child producing a sound correctly for an app does not automatically use that sound correctly during spontaneous conversation. Similarly, an adult may achieve good scores in a naming activity without experiencing equivalent improvement in functional communication.

Nine technologies changing speech and language therapy

1. Automatic speech recognition

Automatic speech recognition, or ASR, converts spoken language into text or interpretable digital information.

Consumers already interact with it through voice assistants, dictation, automatic subtitles, transcription and voice-controlled interfaces.

In speech-language therapy, ASR creates several possibilities. A system might identify whether a target word was produced, generate a transcript, support home practice or allow a person with limited hand function to interact with a computer using speech.

However, most mainstream speech-recognition systems have historically been trained predominantly on typical speech. Performance may decrease with dysarthria, apraxia of speech, severe speech sound disorders, dysphonia, disfluency, atypical rate or prosody, and different accents and dialects.

2. Voice and acoustic analysis

Digital tools can analyze aspects of a recorded voice signal.

Depending on the application and validation, measurements may include fundamental frequency, intensity, duration, speech rate, pauses, selected acoustic characteristics and phonation time.

This can provide useful complementary information in some areas of voice and motor-speech practice.

It also creates a major temptation: more acoustic data does not automatically mean better clinical assessment.

Recording conditions, microphone quality, background noise, software methodology and individual variability can all affect measurements.

3. Artificial intelligence and generative AI

AI can potentially support several areas of speech-language pathology.

Generative AI may assist clinicians in creating vocabulary activities, semantic exercises, reading materials, conversation scenarios, simplified instructions, picture-description prompts, role-play situations, caregiver education and documentation drafts.

For example, an SLP working with a person with aphasia could generate several levels of functional conversation practice around ordering food or using public transport.

A pediatric therapist could create personalized language activities around a child’s preferred interests.

But an AI-generated activity is not automatically a therapeutic intervention. The SLP still needs to decide what function is being targeted, whether vocabulary is developmentally appropriate, whether linguistic complexity is suitable, whether the information is accurate and whether cultural and linguistic characteristics are respected.

4. Augmentative and alternative communication

Augmentative and alternative communication—AAC—is one of the clearest examples of technology directly enabling participation.

AAC may include communication boards, symbols, picture systems, text-based communication, switches, eye-gaze systems, tablet applications and speech-generating devices.

Technology can help a person communicate when natural speech is insufficient for everyday needs.

The goal is not simply to operate the device. The clinically meaningful outcome is whether the person can express needs, make choices, participate in conversations, build relationships and exercise greater autonomy.

5. Virtual reality, social interaction and multi-user rehabilitation

Virtual reality can bring a particularly important dimension to speech and language therapy: communication in context.

Many communication skills are difficult to train naturally inside a traditional therapy room.

A clinician can simulate a conversation using pictures or role play, but it is not the same as placing the person inside a realistic social environment where they need to listen, respond, make decisions and interact.

VR can recreate situations such as ordering food in a café, asking for information in a shop, participating in a classroom, speaking during a job interview, asking for help in a public place, taking part in a group conversation, practicing turn-taking, initiating communication, responding to unexpected questions, interpreting non-verbal social cues and managing communication in noisy or distracting environments.

The therapist can progressively control the complexity of the situation. A simple scenario may involve one virtual character and a limited number of responses. A more advanced scenario may introduce several speakers, background noise, time pressure, unexpected questions or competing social cues.

This can target pragmatic language, social communication, conversation initiation, turn-taking, topic maintenance, perspective taking, functional vocabulary, auditory comprehension, confidence and communication under cognitive load.

Multi-user or dual VR sessions

An especially promising approach is the use of shared virtual environments.

Instead of a patient interacting only with a computer-controlled avatar, the clinician and patient can enter the same virtual environment.

Both can potentially be represented by avatars and communicate in real time.

This allows the speech-language pathologist to model communication directly inside the situation, demonstrate an appropriate response, provide cues without interrupting the scenario, progressively reduce assistance, observe how the person reacts in context and change the environment in real time.

Multi-user environments could also involve two patients or small groups when clinically appropriate.

For example, two children working on social communication might complete a cooperative task together. One may need to ask the other for an object, explain a strategy, negotiate roles or solve a problem collaboratively.

For adults, a shared environment could simulate workplace communication, restaurant interactions, community participation, group decision-making and social conversations.

This is important because certain communication situations are difficult, costly or unsafe to reproduce repeatedly in hospitals, schools or rehabilitation centers.

A therapist cannot easily transform an office into a supermarket, restaurant, airport, classroom or workplace. Virtual environments can make these contexts available repeatedly while the clinician remains able to control difficulty and provide support.

The objective, however, remains transfer to the real world. Success in a virtual café should ultimately support communication in a real café.

6. Tablets, applications and digital therapeutic activities

Tablets can support many activities used in speech and language therapy, including naming, categorization, matching, sequencing, comprehension, vocabulary, phonological awareness, storytelling, sentence construction, memory and social communication.

Digital exercises can be rapidly adapted and may support repeated practice.

However, clinicians need to avoid turning therapy into a sequence of isolated touchscreen tasks. Communication should ultimately generalize to interaction with people and everyday environments.

7. Telepractice

Telepractice allows speech-language services to be delivered remotely through video communication and other connected tools.

Sessions can include assessment components, language activities, articulation practice, fluency intervention, voice therapy, caregiver coaching, AAC support, education and home-program review.

Telepractice can also allow the clinician to observe communication within the person’s home environment and involve family members more directly.

It is not appropriate for every situation. Internet quality, hearing and vision, cognitive ability, caregiver availability, privacy, clinical complexity and local licensing requirements can affect suitability.

8. Automated feedback and speech-training applications

An increasingly important research area involves tools that attempt to provide automatic feedback on speech production.

Potential applications include home practice for speech sound disorders, articulation, pronunciation, motor speech and voice.

The attraction is clear: a person could potentially receive many more practice opportunities between sessions.

But feedback accuracy is critical. Incorrect automatic feedback may reinforce errors or create frustration.

A particularly relevant model is: SLP defines target → technology supports additional practice → SLP reviews progress and adapts intervention.

9. Digital documentation and workflow tools

Technology can also transform work that occurs outside direct therapy.

Speech-language pathologists may use digital systems for scheduling, documentation, transcription, session preparation, report templates, resource creation, outcome tracking and caregiver communication.

Generative AI can potentially reduce some repetitive workload.

But automated documentation introduces substantial responsibility. An AI should never invent test scores, observations, diagnoses, symptoms, treatment responses or progress.

A safer workflow remains: AI-supported draft → SLP review → correction → professional approval.

From digital speech metrics to meaningful communication

Digital metric Possible meaning Important limitation
Correct productions Accuracy during structured practice May not generalize to conversation
Speech rate Fluency or motor-speech information Context changes natural rate
Voice intensity Vocal output Loudness alone does not define healthy voice
Response time Processing or task performance Influenced by interface familiarity
Vocabulary score Performance on a particular task Does not equal functional language
App repetitions Practice volume Does not prove quality
AAC selections Device use Does not capture communicative intent alone
VR interaction success Performance in a simulated situation Must transfer to real-world communication
Session attendance Access and engagement Attendance does not equal therapeutic outcome

The same principle applies across digital rehabilitation: what can be counted is not always what matters most.

For speech and language therapy, the most meaningful question often remains: Can this person communicate more effectively in everyday life?

Five realistic clinical scenarios

Pediatric speech sound intervention

A child working on a target sound practices with the SLP during a session. The therapist determines the appropriate sound, context and cueing strategy.

A digital application may then provide short additional practice sessions at home. Parents receive simple instructions. The app may record practice attempts or provide limited automated feedback.

During the next session, the therapist evaluates whether the target is beginning to appear in less structured speech.

Technology extends practice. It does not choose the clinical target independently.

Aphasia rehabilitation after stroke

A person with aphasia may work on word retrieval, sentence production and functional conversation.

Digital tools might generate personalized images, naming exercises, semantic categories, conversation situations and home-practice materials.

The therapist can progressively move from structured digital exercises toward real conversation.

An improvement in naming accuracy is useful, but the central goal may be something such as successfully making a telephone call, participating in family discussion or ordering independently in a café.

Social communication in a shared virtual environment

A child or adolescent working on social communication may enter a virtual café, shop or classroom with the therapist.

Instead of discussing an imaginary situation from a worksheet, both participants are inside the same scenario.

The therapist may ask the child to initiate a conversation, request information, maintain a topic, respond to another person’s question, interpret a social cue or repair a communication breakdown.

Difficulty can be increased progressively by adding another virtual character, background noise or unexpected events.

The therapist can model appropriate communication directly inside the shared environment. Later, the same skills should be practiced in real-world situations.

AAC for a person with severe motor impairment

A person with limited natural speech may use a tablet-based AAC system or eye-gaze communication device.

The SLP may collaborate with occupational therapists, family members and other professionals to optimize vocabulary organization, access method, positioning, communication strategies and partner training.

Success is measured through communication and participation—not how rapidly the person learns the software interface.

Voice therapy and remote monitoring

A person undergoing voice therapy may complete exercises in clinic and at home.

Digital tools may record selected voice parameters or allow the person to track practice.

The SLP interprets the data alongside auditory-perceptual assessment, symptoms, vocal load and functional voice demands.

An acoustic number should not be treated as a diagnosis in isolation.

Multilingualism, accents and AI: an important challenge

Speech and language technologies are particularly sensitive to linguistic variation.

A system trained predominantly on one language, accent or population may perform less accurately with another.

This has implications for automatic speech recognition, pronunciation scoring, speech-disorder detection, language analysis and automated transcription.

Clinicians need to distinguish between a communication disorder and a linguistic difference.

Technology must not transform accent, dialect, multilingualism or cultural variation into pathology merely because an algorithm was trained on a narrower population.

Can AI reduce the administrative workload of SLPs?

Potentially.

Generative systems may support session-plan drafts, material creation, parent information sheets, documentation templates, language simplification, translations that are subsequently professionally reviewed and summary drafts.

Clinicians may therefore be able to spend less time creating repetitive materials.

But efficiency should never come at the cost of accuracy or confidentiality.

What does the evidence tell us?

It is important not to refer to “technology in speech therapy” as though it were one intervention.

Automatic speech recognition, generative language models, AAC, acoustic analysis, VR and automated articulation feedback solve very different problems.

They therefore require different forms of evidence.

For virtual environments, an additional question is crucial: Does improved performance in a simulation transfer to communication in everyday life?

The key hierarchy is: technical performance → clinical validity → therapeutic usefulness → functional communication → real-world participation.

Nine questions before using a digital speech or language tool

  1. What communication goal are we addressing?
  2. Is this tool appropriate for the person’s language, dialect and communication profile?
  3. Has it been validated for the intended clinical use?
  4. Is automated feedback sufficiently reliable?
  5. Can the person access and understand the interface?
  6. What happens when the technology makes an error?
  7. How will performance inside the tool transfer to real communication?
  8. Does a simulated or virtual activity represent a meaningful real-world goal?
  9. Could a simpler method achieve the same goal?

The future SLP: augmented communication expertise

Technology is likely to continue expanding the amount of information and therapeutic environments available to speech-language pathologists.

Future clinicians may work with acoustic measurements, automatic transcripts, home-practice data, AAC analytics, AI-generated resources, remote-session data, virtual social environments, multi-user rehabilitation sessions and digital outcome measures.

The challenge will not be to maximize the amount of technology used. It will be to decide which technology meaningfully improves communication and participation.

One of the most promising directions may be the combination of therapist expertise with digital environments that reproduce situations unavailable in the clinic.

Instead of practicing only isolated words or scripted dialogues, patients may be able to practice communication inside virtual homes, shops, schools, workplaces and public spaces, with their therapist present inside the same environment.

At Remotion, this principle is particularly relevant to the development of interactive rehabilitation and immersive environments.

A shared virtual environment can give therapists the possibility to create realistic situations, adapt complexity in real time and practice communication skills in contexts that would otherwise be difficult to reproduce inside a rehabilitation center.

The objective is not to replace face-to-face human interaction. It is to extend where, how and under which conditions meaningful communication can be practiced.

Frequently asked questions

Can AI diagnose speech or language disorders?

AI can identify patterns and assist selected assessment processes, but a clinical diagnosis requires validated assessment, professional interpretation and appropriate context.

Can automatic speech recognition assess articulation?

Potentially for specific, validated tasks, but mainstream speech recognition is not automatically a clinical articulation assessment tool.

Can AI replace a speech-language pathologist?

No. AI can support selected tasks, but assessment, treatment planning, communication-partner interaction, clinical reasoning and professional responsibility remain human responsibilities.

Can VR be used for speech and language therapy?

Potentially, yes. Virtual environments can support functional communication, social interaction, conversation practice and contextualized language tasks. Their value depends on the patient’s goals, accessibility and whether skills transfer to real-life communication.

What is a multi-user VR therapy session?

It is a shared virtual environment in which the therapist and patient—or, when appropriate, several participants—can interact in real time. This can allow the clinician to model communication, give cues and practice realistic social situations inside the same virtual scenario.

Is telepractice effective for speech therapy?

Telepractice can be appropriate for many interventions, but effectiveness and suitability depend on the condition, patient, technology, intervention model and environment.

Can AAC prevent someone from developing natural speech?

AAC is designed to support communication when speech alone is insufficient. Selection and implementation should be individualized by appropriately qualified professionals.

Can AI create therapy exercises?

Yes, generative AI can produce draft materials and activity ideas. The SLP must verify their linguistic, developmental and clinical appropriateness before use.

Selected references and further reading

  1. American Speech-Language-Hearing Association. Generative Artificial Intelligence for Clinicians in Audiology and Speech-Language Pathology.
  2. American Speech-Language-Hearing Association. The Role of Artificial Intelligence in Speech Disorders. 2024.
  3. Austin J, Benas K, Caicedo S, et al. Perceptions of Artificial Intelligence and ChatGPT by Speech-Language Pathologists and Students. American Journal of Speech-Language Pathology, 2025.
  4. Baker ZA, Snodgrass TD, Benway NR, Preston JL. A Preliminary Study of Speech-Language Pathologist and Parent Perspectives on Artificial Intelligence Use for Speech Sound Disorder Treatment. Perspectives of the ASHA Special Interest Groups, 2026.
  5. Mallipeddi NV, Mehrotra A, Van Stan JH. Telepractice in the Treatment of Speech and Voice Disorders: What Could the Future Look Like? Perspectives of the ASHA Special Interest Groups, 2023.
  6. Liss J, Berisha V. How Will Artificial Intelligence Reshape Speech-Language Pathology Services and Practice in the Future? ASHA Journals Academy, 2020.
  7. ASHA. Artificial Intelligence: Additional Resources and Guidance.
  8. ASHA. Innovations in Technology — Communication Sciences and Disorders.

← Explore the Digital Rehabilitation Knowledge Hub