Interscalar medicine / long-form research note / January 2025
Predictive Rehabilitation and Risk Forecasting
Rehabilitation is already predictive. Every goal, discharge plan, exercise progression, and follow-up interval contains an expectation about what may happen next. The weakness of the conventional process is not the absence of prognosis, but the fact that prognosis is often episodic, implicit, and difficult to audit. It may be formed during a visit and then remain unchanged while most recovery unfolds elsewhere.
Predictive rehabilitation turns that implicit expectation into a continuously testable clinical process. It combines baseline condition, repeated functional measurements, treatment events, patient-reported experience, device-derived signals, and the social and organizational context of care to estimate a future state at a defined time horizon. The estimate is useful only when it is connected to a proportionate, reviewable action.
In this article, predictive rehabilitation means a clinical and operational approach that:
- defines a functional goal or adverse event in advance;
- estimates the probability, range, or trajectory of that outcome over a stated period;
- shows the observations and assumptions that shaped the estimate;
- links the estimate to a human-governed response;
- records what was done and whether it helped;
- recalibrates when the patient, treatment, or care environment changes.
This is a proposed working definition, not an internationally standardized term. It is deliberately broader than “AI in rehabilitation.” A transparent rule, a statistical model, or a change-point detector may be more appropriate than machine learning. Complexity is justified only when it improves decisions, not merely prediction scores.
What exactly is being predicted?
Several different questions are often compressed into the word “prediction.” They should remain separate because they require different evidence.
Outcome prognosis estimates a future state under a defined care pathway: for example, independent walking at discharge, upper-limb function at three months, return to work, or support needs at home.
Trajectory monitoring asks whether observed recovery is departing from an expected range. It can identify a plateau, sudden decline, excessive pain after progression, or a divergence between capacity in the clinic and performance in daily life.
Safety forecasting estimates the risk of an adverse event such as a fall, avoidable readmission, wound complication, or clinical deterioration. A safety model must specify who receives the alert, how quickly they must respond, and what happens when data are absent.
Engagement and access forecasting estimates the risk of missed sessions, incomplete home practice, device abandonment, or loss to follow-up. These events should not automatically be interpreted as a patient’s lack of motivation. Pain, fatigue, cognitive or language impairment, caregiving load, transport, cost, connectivity, poor instructions, and an unrealistic plan may be the actual causes.
Treatment-response estimation asks which intervention is likely to work better for this patient. This is not the same as predicting an outcome. A variable can predict recovery without being a useful treatment target, and observational associations do not establish that changing the variable will change the outcome. Comparative treatment selection requires causal methods, randomized evidence, or another defensible identification strategy.
Operational forecasting concerns the care system: therapist capacity, delayed equipment, authorization, transport, or a missed transition between inpatient and community services. These risks belong to the provider as much as to the patient.
From records to an event model
A conventional record is organized primarily around encounters. Recovery is a sequence. An event model represents each clinically relevant change with at least:
- the subject and care episode;
- what was planned, observed, reported, or performed;
- the time of occurrence and the time of documentation;
- the source: patient, caregiver, clinician, sensor, device algorithm, or administrative system;
- the functional goal or risk to which the event relates;
- the intervention, dose, and responsible actor;
- the immediate result and later outcome;
- context such as pain, fatigue, cognition, environment, assistance, transport, or access;
- data quality, missingness, provenance, and model version.
The World Health Organization’s International Classification of Functioning, Disability and Health (ICF) is a useful semantic frame because it treats functioning as the interaction of body functions and structures, activities, participation, and environmental factors [O1]. HL7 FHIR resources such as Goal, CarePlan, Observation, and RiskAssessment can carry parts of this structure between systems [O2]. Neither standard, however, proves that a particular measure, model, or clinical inference is valid.
The value of the event model is temporal and explanatory. It can distinguish “the exercise was not performed” from “the exercise was not recorded,” “the patient could not perform it because pain increased,” and “the prescribed equipment never arrived.” It can connect a forecast to the evidence available at that moment and later reveal whether the forecast led to an action. It does not turn temporal order or correlation into causation.
The predictive loop
A safe system is a closed loop rather than a risk score on a dashboard:
goal → observations → forecast → review → action → outcome → recalibration
Each forecast should include the outcome, time horizon, probability or plausible range, uncertainty, principal contributing observations, important missing data, model version, and the action threshold. The output should answer not only “how high is the risk?” but also “what can be done now, by whom, and how will we know whether it worked?”
Thresholds must be derived from consequences and capacity. If a high-risk flag triggers a same-day therapist call, false positives consume scarce clinical time; if the threshold is too high, preventable deterioration may be missed. Discrimination alone—an AUC or accuracy value—is insufficient. Calibration, sensitivity in clinically important groups, false-negative harm, decision-curve or net-benefit analysis, workload, and patient outcomes matter. TRIPOD+AI provides current reporting guidance for prediction models, PROBAST+AI addresses quality and risk of bias, and DECIDE-AI addresses early live clinical evaluation of AI decision support [M1–M3].
What the evidence shows—and does not show
Post-stroke upper-limb prognosis is one of the clearest examples. PREP2 combines early clinical assessment with age, neurological severity, and, for some patients, transcranial magnetic stimulation to predict upper-limb function at three months. The development study included 207 patients [R1]. An earlier implementation study of the PREP approach reported altered therapy content, greater therapist confidence, and a shorter median inpatient stay without worse measured outcomes [R2].
But transportability is not guaranteed. A prospective external evaluation in 143 patients reported 78% overall accuracy while performance was much weaker in intermediate outcome groups and in patients with cognitive disorders [R3]. A 2026 evaluation of predictions made in routine care in 83 patients reported 66% overall accuracy and concluded that the full tool was not yet validated in that setting [R4]. These results do not make prognosis useless; they show why a model must be evaluated by outcome category, setting, subgroup, and intended decision rather than advertised with a single headline number.
Wearable sensors can add information unavailable in a brief examination. In a 2024 inpatient stroke study of 55 people, early inertial-sensor data improved some discharge predictions compared with models based only on patient information and functional assessments, but performance varied by outcome and the sensors added no predictive value for the non-ambulatory subgroup [R5]. This is promising feasibility evidence, not proof of broad clinical benefit.
Recent external validation: performance is not impact
A 2025 prospective multisite study developed a model in 220 patients entering intensive inpatient stroke rehabilitation and tested it externally in 165 more. Its best model estimated the modified Barthel Index at discharge with a median absolute error of 11.5 points in internal testing and 9.2 points externally. Baseline Barthel score, upper-limb motor score, age, and cognitive screening supplied most of the predictive signal [R16]. The practical lesson is not that machine learning has replaced assessment; it is that routinely collected functional, motor, and cognitive measurements remain the foundation against which more complex data must add value.
A July 2026 study used prospective and retrospective data collected at four South Korean rehabilitation institutions to predict ambulation, cognition, and activities of daily living across acute-to-subacute and acute-to-early-chronic windows. Reported internal AUROCs reached 0.897 for ambulation and 0.864 for cognition; external AUROCs reached 0.924 for activities of daily living and 0.892 for cognition. The authors also built a web decision-support prototype [R13]. These are encouraging domain-specific and cross-institutional results, but the paper concludes with predictive feasibility and explicitly leaves prospective clinical evaluation pending. High AUROC therefore supports further testing, not a claim that the interface improves therapy, communication, equity, or outcomes.
Independent validation can also change the apparent certainty of a simpler model. EPOS was tested prospectively in 280 heterogeneous patients who could not walk independently within 72 hours after stroke. For independent gait at three months, the AUROC was 0.74 and the Brier score 0.14. At a probability threshold of 0.5, sensitivity was 0.98 but specificity only 0.32 [R14]. Such a threshold may be appropriate when the main ethical priority is not to miss recovery potential, but it will classify many people as potentially recovering. The validation also used a three-month outcome whereas the original model concerned six-month gait, illustrating why the target, horizon, and action rule must travel together. The authors appropriately call next for impact evaluation on decisions, outcomes, and resources.
Engagement forecasts: useful label, incomplete causal story
An analysis of 2,747 users of a sensorized home movement system used the first three weeks of behavior to forecast persistence into week four. The best model reported an AUC of 0.810, precision-recall AUC of 0.717, and F1 score of 0.703; active days, consistency, selected session length, and repetition rate were the strongest signals [R15]. This is unusually large evidence that platform behavior can identify imminent disengagement. It does not establish why a person stops, whether continued platform use produces clinical benefit, whether the model generalizes to people without this device, or whether any proposed retention intervention works. It also cannot rescue users who leave before enough observation has accumulated. The actionable hypothesis—call, simplify the plan, treat pain, provide access support—still requires a prospective impact study.
The wider literature remains immature. A systematic review of 19 post-stroke machine-learning studies found small samples, few external validations, and substantial heterogeneity in predictors and outcomes [R6]. Another review found that all included models had high or unclear risk of bias [R7]. Digital additions to home exercise may improve short-term adherence, but a review of ten randomized trials found heterogeneous measures, low-to-moderate study quality, and less certain long-term effects [R8].
A 2026 umbrella review helps separate “AI effect” from “technology and dose effect.” Across 32 reviews, its most reproducible signal was activity improvement from technology-assisted post-stroke upper-limb practice, especially robotics with or without virtual reality; superiority for impairment or independence was inconsistent when training exposure was matched and assessment was blinded. It also identified development-to-deployment performance loss and sparse reporting of safety, usability, equity, and cost [R17]. However, the umbrella review deliberately did not perform a formal risk-of-bias assessment or grade certainty, and it included several narrative or conceptual sources. Its findings are therefore a map of recurring signals and gaps, not a pooled proof that AI itself causes better rehabilitation outcomes.
The conclusion is deliberately modest: predictive tools can improve planning in defined settings, but the category “predictive rehabilitation” does not yet have a single evidence base. Evidence is condition-specific, outcome-specific, and workflow-specific.
Practical scenarios
1. Early stroke rehabilitation
At admission, the team records a standardized motor assessment, cognition, language and neglect, pre-stroke function, comorbidities, imaging or neurophysiology when justified, and the patient’s own goals. The model returns a probability distribution for upper-limb capacity or independent gait at a defined date—not a deterministic label. The therapist uses it to discuss realistic goals, select compensatory and restorative priorities, plan discharge, and identify uncertainty requiring reassessment. A poor forecast must not be the sole reason to reduce therapy.
2. Home recovery after orthopedic surgery
The care plan defines pain, range-of-motion, walking, sleep, wound, and participation milestones. Patient reports and validated sensor measures are compared with a personal baseline and an expected trajectory. A combined change—rising pain, falling activity, missed exercises, and an unacknowledged message—can trigger a call. The action may be clinical review, plan simplification, equipment support, or transport assistance. Automatic progression of exercise intensity is a higher-risk function and requires stronger evidence and safeguards than a reminder or clinician-facing flag.
3. Chronic neurological or cardiopulmonary rehabilitation
The objective may be maintenance rather than linear improvement. The model watches for meaningful deviation from an individual baseline: reduced community mobility, increased fatigue or dyspnea, falling participation, repeated cancellation, or caregiver strain. The most useful prediction may concern an access barrier rather than physiology. A transport intervention or a lower-burden schedule may prevent “non-adherence” more effectively than another notification.
4. Service coordination
The same event ledger can forecast gaps in care: an assessment that has not been scheduled, equipment likely to arrive after discharge, or a follow-up that exceeds the safe interval. Here the intervention is directed at the organization, not the patient. This prevents a common ethical error—using contextual data to label vulnerable people as risky while leaving the structural cause untouched.
Data quality before model complexity
Baseline clinical scores are often strong predictors. New sensors should be required to demonstrate incremental value over a simple, clinically credible baseline. For digital measures, technical verification is not enough. The V3 framework separates verification, analytical validation, and clinical validation; a device can measure acceleration accurately yet still fail to measure the clinical concept claimed for the intended population [M4]. V3+ adds usability validation: the intended users, including people with motor, sensory, cognitive, or language impairments, must be able to use the sensor-based technology correctly and sustainably in its actual setting [M5]. A change to hardware, firmware, an algorithm, instructions, or workflow may require new evidence for the affected component. FDA guidance on remotely acquired data similarly emphasizes fitness for purpose, participant usability, data protection, and interpretability in a defined context of use [O6].
Missingness deserves its own model. No step count may mean inactivity, an uncharged device, assisted walking that the device cannot detect, loss of connectivity, skin intolerance, hospitalization, or refusal. Treating all missing data as zero can create both clinical error and inequity. A safe design records why data are missing when possible, exposes uncertainty, and preserves a non-digital route through care.
Models should be validated on data separated by patient and time, then externally in a different service or population. Performance should be reported by clinically relevant subgroups and with confidence intervals. Prospective “silent mode” can reveal data failures and alert volume before predictions influence care. After deployment, calibration, overrides, adverse events, subgroup performance, and data drift must be monitored.
Counterarguments and limits
A prediction can narrow possibility. A pessimistic estimate may lower expectations, therapy dose, or access to services and thereby help produce the poor outcome it predicted. This is a form of self-fulfilling prophecy bias recognized in neuroprognostic research [R9]. Forecasts should therefore express uncertainty, be periodically revised, and never serve as the sole basis for withholding rehabilitation.
Prediction is not explanation. An interpretable feature list can show what the model used, not why the outcome will occur. The well-known debate over proportional recovery after stroke shows how apparently strong predictive relationships can be affected by mathematical coupling, baseline distributions, outcome ceilings, and choices about “recoverer” subgroups [R10].
Context can improve prediction and reproduce injustice. Cost, prior utilization, missed visits, device use, and location may be predictive precisely because access is unequal. A widely used population-health algorithm produced racial bias because health cost was used as a proxy for health need [R11]. In rehabilitation, contextual variables should preferably trigger additional support, not reduced eligibility.
Continuous observation is not neutral. Wearables and home video can make the home an extension of the clinic. Data minimization, informed consent, access controls, retention limits, cybersecurity, and an understandable explanation of secondary use are essential. A patient needs a meaningful way to refuse optional monitoring without losing ordinary care.
Digital systems can exclude the people with greatest need. Access, literacy, dexterity, vision, cognition, language, income, and connectivity affect data generation. Studies of older adults continue to find socioeconomic differences in digital access and use [R12]. “No data” must not become “low priority.”
Alerts can move rather than solve work. A model with good retrospective accuracy may fail when the team has no capacity to act, when responsibility is unclear, or when repeated low-value alerts are ignored. Clinical utility requires workflow and impact evaluation, not only technical validation.
Patents are not evidence. Patent documents show that ideas such as adherence prediction, automatic protocol modification, remote-compliance optimization, wearable recovery-time and progress estimation, and post-operative functional tracking have been claimed or disclosed [P1–P6]. They neither demonstrate safety or effectiveness nor establish freedom to operate. A granted patent is a legal right over claims, not a clinical endorsement; patent scope and legal status require specialist review.
A minimum implementation standard
Before deployment, a predictive rehabilitation service should be able to answer:
- What exact outcome and time horizon are predicted?
- What decision changes when the threshold is crossed?
- Is that action effective, available, and proportionate to risk?
- What is the simple clinical baseline, and does the model add value?
- Were inputs and digital measures validated for this population and context?
- How are missing data, uncertainty, and out-of-distribution patients handled?
- Has calibration and clinical utility been externally and prospectively evaluated?
- How does performance differ by age, sex, disability, language, socioeconomic position, site, and device?
- Can clinicians and patients inspect, question, and override the output?
- Who monitors drift, safety events, workload, and unintended denial of care?
If the function meets the definition of medical-device software, lifecycle risk management and clinical evaluation become formal obligations. ISO 14971:2019 provides a medical-device risk-management framework; IMDRF N41 describes clinical evaluation for Software as a Medical Device; FDA guidance distinguishes some clinical decision-support functions from regulated device functions according to intended use and how recommendations are presented [O3–O5]. IMDRF N88 adds ten good-machine-learning-practice principles spanning the total product lifecycle [O7]. Depending on intended use and jurisdiction, IEC 62304 may inform software lifecycle processes, IEC 62366-1 usability engineering related to safety and use error, and IEC 82304-1 product safety and security for health software on general computing platforms [O8–O10]. These documents are not interchangeable certifications, and their current editions and applicability must be checked for the particular product. Classification is jurisdiction- and use-specific and should not be inferred from the presence or absence of “AI.”
Conclusion
The strongest form of predictive rehabilitation is not a model that assigns patients to “good” and “poor” futures. It is a learning care loop that makes goals explicit, detects meaningful deviation early, directs attention to a feasible response, and measures whether that response improved function, participation, safety, or access.
Its central promise is contextual: to see recovery not as a sequence of isolated visits, but as an evolving interaction among biology, function, behavior, care delivery, and environment. Its central discipline is equally important: every forecast must be calibrated, contestable, linked to action, and prevented from becoming a justification for less care.
Интерскалярная медицина / полноформатная исследовательская статья / январь 2025
Предиктивная реабилитация и прогноз риска
Реабилитация уже содержит прогноз. Цель, план выписки, изменение нагрузки и интервал между контрольными визитами всегда основаны на ожидании того, что произойдет дальше. Слабость традиционного процесса не в отсутствии прогноза, а в том, что он часто остается эпизодическим, неявным и плохо проверяемым. Он формируется во время визита и может не меняться, хотя основная часть восстановления происходит за пределами кабинета.
Предиктивная реабилитация превращает такое ожидание в непрерывно проверяемый клинический процесс. Она объединяет исходное состояние, повторные функциональные измерения, события лечения, сообщения пациента, сигналы устройств, а также социальный и организационный контекст, чтобы оценить будущее состояние на заданном временном горизонте. Оценка имеет смысл только тогда, когда она связана с соразмерным и проверяемым действием.
В этой статье предиктивная реабилитация — это клинический и организационный подход, который:
- заранее определяет функциональную цель или нежелательное событие;
- оценивает вероятность, диапазон или траекторию исхода за указанный период;
- показывает наблюдения и допущения, повлиявшие на оценку;
- связывает оценку с действием под контролем человека;
- фиксирует, что было сделано и помогло ли это;
- пересматривает прогноз при изменении пациента, лечения или среды оказания помощи.
Это рабочее определение, а не международно стандартизированный термин. Оно намеренно шире понятия «ИИ в реабилитации». Прозрачное правило, статистическая модель или алгоритм обнаружения изменения траектории могут быть уместнее машинного обучения. Сложность оправдана только тогда, когда она улучшает решения, а не просто метрики прогноза.
Что именно прогнозируется?
Словом «прогноз» часто называют разные задачи. Их необходимо разделять, поскольку для них нужны разные виды доказательств.
Прогноз функционального исхода оценивает будущее состояние при определенном маршруте помощи: например, самостоятельную ходьбу к выписке, функцию верхней конечности через три месяца, возвращение к работе или потребность в поддержке дома.
Мониторинг траектории определяет, выходит ли фактическое восстановление за ожидаемый диапазон. Так можно выявить плато, внезапное ухудшение, чрезмерную боль после увеличения нагрузки или расхождение между возможностями в клинике и повседневной активностью.
Прогноз безопасности оценивает риск падения, предотвратимой повторной госпитализации, раневого осложнения или клинического ухудшения. Для такого прогноза необходимо заранее определить, кто получит сигнал, как быстро должен отреагировать и что делать при отсутствии данных.
Прогноз вовлеченности и доступности помощи оценивает риск пропуска занятий, невыполнения домашней программы, отказа от устройства или потери контакта с пациентом. Эти события нельзя автоматически трактовать как недостаток мотивации. Их причинами могут быть боль, утомляемость, когнитивные или речевые нарушения, нагрузка на семью, транспорт, стоимость, связь, непонятная инструкция или нереалистичный план.
Оценка ответа на лечение отвечает на вопрос, какое вмешательство вероятнее поможет конкретному пациенту. Это не то же самое, что прогноз исхода. Признак может предсказывать восстановление, не являясь мишенью лечения; наблюдаемая связь не доказывает, что изменение признака изменит исход. Для выбора между вмешательствами нужны причинные методы, рандомизированные данные или иная обоснованная стратегия идентификации эффекта.
Организационный прогноз относится к самой системе помощи: доступности специалиста, задержке оборудования, согласованию оплаты, транспорту или разрыву между стационарным и амбулаторным этапами. Ответственность за такие риски лежит на организации не меньше, чем на пациенте.
От медицинской записи к событийной модели
Обычная медицинская запись организована прежде всего вокруг контактов с системой. Восстановление представляет собой последовательность. Событийная модель описывает каждое значимое изменение как минимум через:
- пациента и эпизод помощи;
- планируемое, наблюдаемое, сообщенное или выполненное действие;
- время события и время его документирования;
- источник: пациент, родственник, специалист, сенсор, алгоритм устройства или административная система;
- функциональную цель или риск, к которому относится событие;
- вмешательство, дозу и ответственного исполнителя;
- непосредственный результат и отдаленный исход;
- контекст: боль, утомляемость, когнитивные функции, среда, помощь другого человека, транспорт или доступ;
- качество данных, причины пропусков, происхождение данных и версию модели.
Международная классификация функционирования, ограничений жизнедеятельности и здоровья ВОЗ (МКФ) дает полезную семантическую основу: функционирование рассматривается как взаимодействие функций и структур организма, активности, участия и факторов среды [O1]. Ресурсы HL7 FHIR Goal, CarePlan, Observation и RiskAssessment позволяют передавать части этой структуры между системами [O2]. Однако ни один стандарт сам по себе не подтверждает валидность конкретной шкалы, модели или клинического вывода.
Смысл событийной модели — во временной структуре и возможности объяснить процесс. Она различает ситуации «упражнение не выполнено», «выполнение не зарегистрировано», «пациент не смог выполнить упражнение из-за усиления боли» и «назначенное оборудование не было доставлено». Она связывает прогноз с информацией, доступной в момент его формирования, и позволяет позднее проверить, последовало ли действие. При этом временная последовательность и корреляция не становятся причинностью автоматически.
Предиктивный цикл
Безопасная система — это замкнутый цикл, а не показатель риска на панели:
цель → наблюдения → прогноз → проверка → действие → исход → перекалибровка
Каждый прогноз должен содержать исход, временной горизонт, вероятность или правдоподобный диапазон, неопределенность, основные повлиявшие наблюдения, значимые пропуски данных, версию модели и порог действия. Результат должен отвечать не только на вопрос «насколько высок риск?», но и на вопросы «что можно сделать сейчас, кто это сделает и как понять, помогло ли действие?»
Порог определяется последствиями и реальными ресурсами. Если высокий риск вызывает звонок специалиста в тот же день, ложноположительные сигналы расходуют дефицитное время; если порог слишком высок, можно пропустить предотвратимое ухудшение. Одной дискриминации — AUC или точности — недостаточно. Важны калибровка, чувствительность в клинически значимых группах, вред ложноотрицательных результатов, анализ чистой пользы, нагрузка на команду и исходы для пациента. TRIPOD+AI задает актуальные требования к отчетности по моделям прогноза, PROBAST+AI — к оценке качества и риска систематической ошибки, DECIDE-AI — к ранней клинической проверке систем поддержки решений в реальной работе [M1–M3].
Что показывают и чего не показывают исследования
Один из наиболее разработанных примеров — прогноз функции верхней конечности после инсульта. PREP2 объединяет раннюю клиническую оценку с возрастом, тяжестью неврологического дефицита и, для части пациентов, транскраниальной магнитной стимуляцией, чтобы предсказать функцию руки через три месяца. В исследование разработки вошли 207 пациентов [R1]. Более раннее исследование внедрения подхода PREP показало изменение содержания терапии, рост уверенности специалистов и сокращение медианы пребывания в стационаре без ухудшения измеренных исходов [R2].
Но переносимость модели в другую среду не гарантирована. Проспективная внешняя оценка на 143 пациентах показала общую точность 78%, при этом результат был существенно слабее для промежуточных категорий исхода и пациентов с когнитивными нарушениями [R3]. В оценке 2026 года, где прогноз PREP2 формировался специалистами в обычной практике для 83 пациентов, общая точность составила 66%; авторы заключили, что полный инструмент в этой среде пока не валидирован [R4]. Это не означает, что прогноз бесполезен. Это означает, что модель нужно оценивать по категориям исхода, месту применения, подгруппам и решению, которое она поддерживает, а не рекламировать одной цифрой.
Носимые сенсоры способны добавить информацию, недоступную при кратком осмотре. В исследовании 2024 года с участием 55 пациентов стационара ранние данные инерциальных сенсоров улучшили часть прогнозов к выписке по сравнению с моделями, использовавшими только сведения о пациенте и функциональные шкалы. Однако результат зависел от исхода, а в подгруппе неходячих пациентов сенсоры не дали дополнительной прогностической ценности [R5]. Это обнадеживающая проверка осуществимости, но не доказательство широкой клинической пользы.
Новая внешняя валидация: качество прогноза не равно влиянию на практику
В проспективном многоцентровом исследовании 2025 года модель разработали на данных 220 пациентов, поступивших на интенсивную стационарную реабилитацию после инсульта, и внешне проверили еще на 165 пациентах. Лучшая модель оценивала модифицированный индекс Бартел к выписке с медианной абсолютной ошибкой 11,5 балла при внутренней и 9,2 балла при внешней проверке. Основной вклад внесли исходный индекс Бартел, моторная оценка верхней конечности, возраст и когнитивный скрининг [R16]. Практический вывод не в том, что машинное обучение заменило обследование. Наоборот, обычные функциональные, моторные и когнитивные показатели остаются основой, по отношению к которой более сложные данные должны доказывать дополнительную ценность.
В работе июля 2026 года использованы проспективные и ретроспективные данные четырех реабилитационных организаций Южной Кореи для прогноза ходьбы, когнитивной функции и повседневной активности на интервалах от острой до подострой и ранней хронической фазы. Максимальные AUROC при внутренней проверке составили 0,897 для ходьбы и 0,864 для когнитивной функции, при внешней — 0,924 для повседневной активности и 0,892 для когнитивной функции. Авторы также создали веб-прототип поддержки решений [R13]. Это обнадеживающие результаты по отдельным функциональным доменам и между организациями, но вывод статьи ограничен прогностической осуществимостью, а проспективная клиническая оценка еще предстоит. Высокая AUROC оправдывает следующую проверку, но не доказывает, что интерфейс улучшает терапию, коммуникацию, справедливость или исходы.
Независимая проверка способна снизить кажущуюся определенность и более простой модели. EPOS проспективно оценили у 280 разнородных пациентов, которые не могли самостоятельно ходить в первые 72 часа после инсульта. Для самостоятельной ходьбы через три месяца AUROC составила 0,74, показатель Брайера — 0,14. При пороге вероятности 0,5 чувствительность достигла 0,98, но специфичность была лишь 0,32 [R14]. Такой порог уместен, если главный этический приоритет — не пропустить потенциал восстановления, но тогда модель отнесет к потенциально восстанавливающимся многих пациентов. Кроме того, валидация использовала трехмесячный исход, тогда как исходная модель касалась ходьбы через шесть месяцев: целевой исход, горизонт и правило действия должны переноситься вместе. Авторы обоснованно предлагают следующим шагом оценить влияние модели на решения, исходы и ресурсы.
Прогноз вовлеченности: полезная метка, но неполное причинное объяснение
Анализ 2747 пользователей сенсорной системы домашних двигательных упражнений использовал поведение за первые три недели, чтобы спрогнозировать продолжение занятий на четвертой неделе. Лучшая модель показала AUC 0,810, AUC кривой точность–полнота 0,717 и F1 0,703; наиболее значимыми признаками стали число активных дней, регулярность, выбранная длительность занятия и темп повторений [R15]. Это редкий по размеру массив данных, показывающий, что поведение на платформе позволяет выявлять близкое прекращение занятий. Однако он не объясняет причину прекращения, не доказывает клиническую пользу продолжения работы именно с этой платформой, переносимость на людей без этого устройства или эффективность мер по удержанию. Модель также не поможет тем, кто уйдет до накопления достаточного периода наблюдения. Рабочая гипотеза — позвонить, упростить план, лечить боль или устранить барьер доступа — по-прежнему требует проспективной проверки влияния.
Общая литература остается незрелой. Систематический обзор 19 работ по машинному обучению после инсульта выявил небольшие выборки, редкую внешнюю валидацию и высокую неоднородность признаков и исходов [R6]. В другом обзоре все включенные модели имели высокий или неопределенный риск систематической ошибки [R7]. Цифровые дополнения к домашним упражнениям, вероятно, повышают краткосрочную приверженность, но обзор десяти рандомизированных исследований обнаружил неоднородные методы измерения, низкое или среднее качество работ и меньшую определенность долгосрочного эффекта [R8].
Зонтичный обзор 2026 года помогает отделить «эффект ИИ» от эффекта технологии и дозы занятий. Среди 32 обзоров наиболее воспроизводимым сигналом было улучшение активности при технологически поддерживаемой тренировке верхней конечности после инсульта, особенно с роботами и виртуальной реальностью; превосходство по нарушениям функций или самостоятельности становилось непоследовательным при сопоставимой тренировочной нагрузке и ослепленной оценке. Авторы также выявили падение качества от разработки к реальному внедрению и редкую отчетность по безопасности, удобству, справедливости и стоимости [R17]. Однако формальная оценка риска систематической ошибки и уровня уверенности в доказательствах не проводилась, а среди источников были нарративные и концептуальные работы. Поэтому это карта повторяющихся сигналов и пробелов, а не объединенное доказательство того, что именно ИИ улучшает исходы реабилитации.
Вывод намеренно сдержанный: предиктивные инструменты могут улучшать планирование в конкретной среде, но у всей категории «предиктивная реабилитация» пока нет единой доказательной базы. Доказательства зависят от заболевания, исхода и рабочего процесса.
Практические сценарии
1. Ранний этап реабилитации после инсульта
При поступлении команда фиксирует стандартизированную моторную оценку, когнитивные и речевые функции, игнорирование пространства, состояние до инсульта, сопутствующие заболевания, а при обоснованной необходимости — нейровизуализацию или нейрофизиологию, а также личные цели пациента. Модель возвращает распределение вероятностей функции руки или самостоятельной ходьбы к определенной дате, а не окончательный ярлык. Специалист использует результат для обсуждения реалистичных целей, выбора восстановительных и компенсаторных приоритетов, планирования выписки и определения момента переоценки. Неблагоприятный прогноз не должен быть единственным основанием для уменьшения терапии.
2. Домашнее восстановление после ортопедической операции
План определяет контрольные точки для боли, объема движения, ходьбы, сна, состояния раны и участия в повседневной жизни. Сообщения пациента и валидированные сенсорные показатели сравниваются с персональным исходным уровнем и ожидаемой траекторией. Сочетание изменений — нарастание боли, снижение активности, пропуски упражнений и сообщение без ответа — может вызвать звонок. Действием может быть клиническая проверка, упрощение плана, помощь с оборудованием или транспортом. Автоматическое увеличение нагрузки несет больший риск и требует более сильных доказательств и ограничений, чем напоминание или сигнал специалисту.
3. Хроническая неврологическая или кардиореспираторная реабилитация
Целью может быть поддержание функции, а не линейное улучшение. Модель отслеживает значимое отклонение от персонального уровня: снижение мобильности вне дома, рост утомляемости или одышки, уменьшение участия, повторные отмены занятий или нагрузку на родственника. Наиболее полезный прогноз может касаться не физиологии, а доступности помощи. Транспорт или менее обременительное расписание иногда предотвращают «неприверженность» эффективнее очередного уведомления.
4. Координация помощи
Тот же журнал событий может прогнозировать разрывы: обследование не назначено, оборудование прибудет после выписки, интервал до контроля превышает безопасный. Здесь вмешательство направлено на организацию, а не на пациента. Это предотвращает распространенную этическую ошибку, когда контекстные данные используются, чтобы назвать уязвимого человека «рискованным», но структурная причина остается без изменения.
Сначала качество данных, затем сложность модели
Исходные клинические шкалы нередко являются сильными предикторами. Новый сенсор должен доказать дополнительную ценность по сравнению с простой и клинически правдоподобной базовой моделью. Для цифрового показателя технической точности недостаточно. Схема V3 разделяет техническую верификацию, аналитическую и клиническую валидацию: устройство может точно измерять ускорение и при этом не измерять заявленное клиническое понятие в целевой популяции [M4]. V3+ добавляет валидацию удобства использования: целевые пользователи, включая людей с двигательными, сенсорными, когнитивными или речевыми нарушениями, должны правильно и устойчиво пользоваться сенсорной технологией в реальной среде [M5]. Изменение аппаратной части, прошивки, алгоритма, инструкции или рабочего процесса может потребовать повторного подтверждения затронутого компонента. Руководство FDA по дистанционному сбору данных также подчеркивает соответствие назначению, удобство для участника, защиту данных и интерпретируемость в заданном контексте применения [O6].
Пропуски данных требуют отдельной модели. Отсутствие числа шагов может означать неподвижность, разряженное устройство, ходьбу с помощью другого человека, которую сенсор не распознает, отсутствие связи, раздражение кожи, госпитализацию или отказ от наблюдения. Если любой пропуск считать нулем, возникают клинические ошибки и неравенство. Безопасная система по возможности фиксирует причину пропуска, показывает неопределенность и сохраняет нецифровой маршрут помощи.
Модель следует валидировать с разделением данных по пациентам и времени, а затем — внешне, в другой службе или популяции. Результаты публикуются для клинически значимых подгрупп и с доверительными интервалами. Проспективный «тихий режим», когда прогноз еще не влияет на помощь, выявляет сбои данных и реальный объем сигналов. После внедрения необходимо отслеживать калибровку, отмены и исправления решений, нежелательные события, качество по подгруппам и дрейф данных.
Контраргументы и ограничения
Прогноз может сузить поле возможного. Пессимистическая оценка способна снизить ожидания, дозу терапии или доступ к услугам и тем самым помочь создать предсказанный неблагоприятный исход. В исследованиях нейропрогноза это называют смещением самоисполняющегося пророчества [R9]. Поэтому прогноз должен показывать неопределенность, регулярно пересматриваться и никогда не быть единственным основанием для отказа в реабилитации.
Прогноз не равен объяснению. Перечень значимых признаков показывает, чем пользовалась модель, но не доказывает причину исхода. Дискуссия о правиле пропорционального восстановления после инсульта демонстрирует, как сильные на вид связи зависят от математического сопряжения, распределения исходных оценок, потолочных эффектов и способа выделения «восстанавливающихся» пациентов [R10].
Контекст улучшает прогноз и может воспроизводить несправедливость. Стоимость помощи, предыдущие обращения, пропуски визитов, использование устройств и место проживания могут быть прогностичны именно потому, что доступ распределен неравномерно. Известный алгоритм управления здоровьем населения создал расовое смещение, поскольку финансовые затраты использовались как заместитель потребности в помощи [R11]. В реабилитации контекстные признаки предпочтительно использовать для предоставления дополнительной поддержки, а не для ограничения допуска.
Непрерывное наблюдение не нейтрально. Носимые устройства и домашнее видео превращают дом в продолжение клиники. Необходимы минимизация данных, информированное согласие, разграничение доступа, сроки хранения, кибербезопасность и понятное описание вторичного использования. Пациент должен иметь реальную возможность отказаться от необязательного мониторинга, не потеряв обычную помощь.
Цифровая система может исключить тех, кому помощь особенно нужна. Доступ к устройству, грамотность, моторика, зрение, когнитивные и речевые функции, доход и связь влияют на образование данных. Исследования пожилых людей по-прежнему выявляют социально-экономические различия в цифровом доступе и использовании [R12]. «Нет данных» не должно означать «низкий приоритет».
Сигналы могут переносить, а не решать работу. Ретроспективно точная модель провалится, если у команды нет ресурса на действие, ответственность не определена или многочисленные малополезные уведомления игнорируются. Клиническая полезность требует проверки рабочего процесса и влияния на исходы, а не только технической валидации.
Патент не является доказательством. В патентах раскрыты или заявлены прогнозирование приверженности, автоматическое изменение протокола, оптимизация дистанционного соблюдения программы, оценка срока и хода восстановления по носимым устройствам и послеоперационное отслеживание функции [P1–P6]. Это не доказывает безопасность или эффективность и не подтверждает свободу использования технологии. Выданный патент — юридическое право на формулу изобретения, а не клиническое одобрение; объем прав и юридический статус требуют отдельной профессиональной проверки.
Минимальный стандарт внедрения
До запуска служба предиктивной реабилитации должна ответить на десять вопросов:
- Какой именно исход и на каком горизонте прогнозируется?
- Какое решение меняется при пересечении порога?
- Доступно ли это действие, доказана ли его польза и соразмерно ли оно риску?
- Какова простая клиническая базовая модель и дает ли новая модель дополнительную ценность?
- Валидированы ли исходные показатели и цифровые измерения для этой популяции и среды?
- Как обрабатываются пропуски, неопределенность и пациенты вне распределения обучающих данных?
- Проверены ли калибровка и клиническая полезность внешне и проспективно?
- Как различается качество по возрасту, полу, инвалидности, языку, социально-экономическому положению, организации и устройству?
- Могут ли специалист и пациент проверить, оспорить и отменить результат?
- Кто отслеживает дрейф, события безопасности, нагрузку и непреднамеренное ограничение помощи?
Если функция соответствует определению программного медицинского изделия, управление риском жизненного цикла и клиническая оценка становятся формальными обязанностями. ISO 14971:2019 задает рамку управления риском медицинских изделий; IMDRF N41 описывает клиническую оценку Software as a Medical Device; руководство FDA разделяет некоторые функции клинической поддержки решений и регулируемые функции медицинского изделия в зависимости от назначения и способа представления рекомендаций [O3–O5]. IMDRF N88 добавляет десять принципов надлежащей практики машинного обучения на протяжении всего жизненного цикла продукта [O7]. В зависимости от назначения и юрисдикции IEC 62304 может задавать процессы жизненного цикла ПО, IEC 62366-1 — проектирование удобства использования в части безопасности и ошибок применения, IEC 82304-1 — безопасность и защищенность программного продукта здравоохранения на универсальной вычислительной платформе [O8–O10]. Эти документы не являются взаимозаменяемыми сертификатами; для конкретного продукта нужно проверить применимость и актуальные редакции. Классификация зависит от юрисдикции и конкретного применения, а не просто от наличия или отсутствия слова «ИИ».
Вывод
Сильная предиктивная реабилитация — не модель, которая делит людей на пациентов с «хорошим» и «плохим» будущим. Это обучающийся контур помощи, который делает цели явными, рано замечает значимое отклонение, направляет внимание к выполнимому действию и проверяет, улучшило ли это действие функцию, участие, безопасность или доступность помощи.
Главное обещание подхода — контекст: видеть восстановление не как набор изолированных визитов, а как меняющееся взаимодействие биологии, функции, поведения, организации помощи и среды. Не менее важна дисциплина: каждый прогноз должен быть калиброванным, оспоримым, связанным с действием и защищенным от превращения в оправдание меньшего объема помощи.
References / Список литературы
The literature list is separate from the article text. All URLs below were checked on 28–29 July 2026; DOI, PMID, document codes, and patent publication numbers are included where available. / Список литературы отделен от основного текста. Все ссылки проверены 28–29 июля 2026 года; при наличии указаны DOI, PMID, коды официальных документов и номера патентных публикаций.
Peer-reviewed rehabilitation evidence / Рецензируемые исследования реабилитации
- [R1] Stinear CM et al. PREP2: A biomarker-based algorithm for predicting upper limb function after stroke. Annals of Clinical and Translational Neurology (2017). DOI 10.1002/acn3.488; PMID 29159193. Development cohort of 207 patients.
- [R2] Stinear CM et al. Predicting Recovery Potential for Individual Stroke Patients Increases Rehabilitation Efficiency. Stroke (2017). DOI 10.1161/STROKEAHA.116.015790; PMID 28280137. Clinical implementation study; not a blinded randomized impact trial.
- [R3] Millot S et al. Prediction of Upper Limb Motor Recovery by the PREP2 Algorithm in a Nonselected Population: External Validation and Influence of Cognitive Syndromes. Neurorehabilitation and Neural Repair (2024). DOI 10.1177/15459683241270056; PMID 39162251.
- [R4] Jordan H, Norrie O, Stinear CM. The Accuracy of the PREP2 Prediction Tool for Upper Limb Outcomes After Stroke as Part of Routine Clinical Care. Neurorehabilitation and Neural Repair (2026). DOI 10.1177/15459683251412283; PMID 41574464.
- [R5] O’Brien MK et al. Early Prediction of Poststroke Rehabilitation Outcomes Using Wearable Sensors. Physical Therapy (2024). DOI 10.1093/ptj/pzad183; PMID 38169444.
- [R6] Campagnini S et al. Machine learning methods for functional recovery prediction and prognosis in post-stroke rehabilitation: a systematic review. Journal of NeuroEngineering and Rehabilitation (2022). DOI 10.1186/s12984-022-01032-4; PMID 35659246.
- [R7] Zu W et al. Machine learning in predicting outcomes for stroke patients following rehabilitation treatment: A systematic review. PLOS ONE (2023). DOI 10.1371/journal.pone.0287308; PMID 37379289.
- [R8] Lang S et al. Do digital interventions increase adherence to home exercise rehabilitation? A systematic review of randomised controlled trials. Archives of Physiotherapy (2022). DOI 10.1186/s40945-022-00148-z; PMID 36184611.
- [R9] Mainali S et al. Do Neuroprognostic Studies Account for Self-Fulfilling Prophecy Bias in Their Methodology? The SPIN Protocol for a Systematic Review. Critical Care Explorations (2023). DOI 10.1097/CCE.0000000000000943; PMID 37396931. Used to define the bias mechanism; not rehabilitation-effect evidence.
- [R10] Kundert R et al. What the Proportional Recovery Rule Is (and Is Not): Methodological and Statistical Considerations. Neurorehabilitation and Neural Repair (2019). DOI 10.1177/1545968319872996; PMID 31524062.
- [R11] Obermeyer Z et al. Dissecting racial bias in an algorithm used to manage the health of populations. Science (2019). DOI 10.1126/science.aax2342; PMID 31649194. General health-algorithm evidence; not rehabilitation-specific.
- [R12] Yang R, Gao S, Jiang Y. Digital divide as a determinant of health in U.S. older adults: prevalence, trends, and risk factors. BMC Geriatrics (2024). DOI 10.1186/s12877-024-05612-y; PMID 39709341.
- [R13] Kim D-Y et al. Domain-specific functional outcome prediction in stroke rehabilitation: A multicenter artificial intelligence study. International Journal of Medical Informatics (2026), 220:106607. DOI 10.1016/j.ijmedinf.2026.106607; PMID 42485958. Multicenter model and decision-support prototype; prospective clinical impact evaluation remains pending.
- [R14] Vinzens L, Betschart M, Veerbeek JM. External Validation of the EPOS Prediction Model for Independent Gait After Stroke. Neurorehabilitation and Neural Repair (2026), online ahead of print. DOI 10.1177/15459683261425934; PMID 42136245. Prospective external validation, n=280; trial registrations NCT06438770 and NCT05039047.
- [R15] Kim SJ et al. Forecasting Dropout in Home-Based Movement Rehabilitation After Stroke With Sensors and Machine Learning. IEEE Transactions on Neural Systems and Rehabilitation Engineering (2026), 34:2438–2446. DOI 10.1109/TNSRE.2026.3692207; PMID 42113670. Sensorized-system behavior study, n=2,747; predicts platform persistence rather than a clinical outcome.
- [R16] Campagnini S et al. Prediction of the functional outcome of intensive inpatient rehabilitation after stroke using machine learning methods. Scientific Reports (2025), 15:16083. DOI 10.1038/s41598-025-00781-1; PMID 40341247; PMCID PMC12062331. Prospective multisite development and external validation, total n=385.
- [R17] Abdalla N et al. Artificial intelligence in rehabilitation: a review of clinical effectiveness, real-world performance, safety, and equity across modalities and settings. Frontiers in Digital Health (2026), 8:1737957. DOI 10.3389/fdgth.2026.1737957; PMID 41929610; PMCID PMC13040452; publisher full text. Umbrella review of 32 reviews; no formal risk-of-bias assessment or certainty grading was performed.
Measurement and model-evaluation methods / Методология измерений и оценки моделей
- [M1] Collins GS et al. TRIPOD+AI statement: updated guidance for reporting clinical prediction models that use regression or machine learning methods. BMJ (2024). DOI 10.1136/bmj-2023-078378; PMID 38626948. Supersedes the 2015 TRIPOD checklist.
- [M2] Moons KGM et al. PROBAST+AI: an updated quality, risk of bias, and applicability assessment tool for prediction models using regression or artificial intelligence methods. BMJ (2025). DOI 10.1136/bmj-2024-082505; PMID 40127903; publisher full text.
- [M3] Vasey B et al. Reporting guideline for the early stage clinical evaluation of decision support systems driven by artificial intelligence: DECIDE-AI. BMJ (2022). DOI 10.1136/bmj-2022-070904; PMID 35584845.
- [M4] Goldsack JC et al. Verification, analytical validation, and clinical validation (V3): the foundation of determining fit-for-purpose for Biometric Monitoring Technologies. npj Digital Medicine (2020). DOI 10.1038/s41746-020-0260-4; PMID 32337371.
- [M5] Bakker JP et al. V3+ extends the V3 framework to ensure user-centricity and scalability of sensor-based digital health technologies. npj Digital Medicine (2025), 8:51. DOI 10.1038/s41746-024-01322-2; PMID 39856145; PMCID PMC11760348. Adds usability validation to V3; the article reports extensive author relationships with digital-health organizations and companies.
Official classifications, standards, and regulatory guidance / Официальные классификации, стандарты и руководства
- [O1] World Health Organization. International Classification of Functioning, Disability and Health (ICF). Official WHO classification page. Endorsed in WHA 54.21 (2001).
- [O2] HL7 FHIR R4. Official resources: Goal, CarePlan, Observation, and RiskAssessment.
- [O3] ISO 14971:2019. Medical devices—Application of risk management to medical devices. Official ISO record. The full standard is not open access; the official bibliographic page is.
- [O4] International Medical Device Regulators Forum. Software as a Medical Device (SaMD): Clinical Evaluation. IMDRF/SaMD WG/N41FINAL:2017. Official IMDRF document page.
- [O5] U.S. Food and Drug Administration. Clinical Decision Support Software: Guidance for Industry and FDA Staff. January 2026, docket FDA-2017-D-6569. Official FDA guidance page.
- [O6] U.S. Food and Drug Administration. Digital Health Technologies for Remote Data Acquisition in Clinical Investigations. December 2023, docket FDA-2021-D-1128. Official FDA guidance page.
- [O7] International Medical Device Regulators Forum. Good machine learning practice for medical device development: Guiding principles. IMDRF/AIML WG/N88 FINAL:2025, 29 January 2025. Official IMDRF document page.
- [O8] IEC 62304:2006. Medical device software—Software life cycle processes; Amendment 1:2015. Official ISO/IEC bibliographic record. The record described the 2006 edition as current but under systematic review on 29 July 2026; the full standard is not open access.
- [O9] IEC 62366-1:2015. Medical devices—Part 1: Application of usability engineering to medical devices; Amendment 1:2020. Official ISO/IEC bibliographic record. The record listed the edition as published and confirmed when checked; the full standard is not open access.
- [O10] IEC 82304-1:2016. Health software—Part 1: General requirements for product safety. Official ISO/IEC bibliographic record. The record described the standard as published and under systematic review on 29 July 2026; the full standard is not open access.
Patent documents: thematic prior art, not clinical evidence / Патентные документы: уровень техники, а не клинические доказательства
- [P1] US20180096111A1, “Predictive telerehabilitation technology and user interface.” Published 5 April 2018; predicts future therapy adherence and describes protocol modification. Patent document. The aggregator listed the U.S. application as abandoned when checked; legal status must be independently verified before relying on it.
- [P2] US9129054B2, “Systems and methods for surgical and interventional planning, support, post-operative follow-up, and functional recovery tracking.” Granted 8 September 2015; includes monitoring treatment efficacy and suggesting changes to post-operative plans. Patent document.
- [P3] WO2023044052A1, “Predicting subjective recovery from acute events using consumer wearables.” Published 23 March 2023; describes prediction of recovery time from pre- and post-event wearable data. Patent document.
- [P4] US20230395228A1, “Physical therapy imaging and prediction.” Published 7 December 2023; describes risk-based creation or modification of digital home exercise plans using camera, sensor, and clinical inputs. Patent document. The aggregator listed the application as pending when checked; legal status must be independently verified.
- [P5] US11328807B2, “System and method for using artificial intelligence in telemedicine-enabled hardware to optimize rehabilitative routines capable of enabling remote rehabilitative compliance.” Granted 10 May 2022; describes machine-learning estimates of exercise-regimen benefit and contextual factors for remote rehabilitation. Patent document. Google Patents displayed ROM Technologies, Inc. as assignee and “Active” as status when checked; neither entry is a legal opinion.
- [P6] US12011291B2, “Non-invasive wearable biomechanical and physiology monitor for injury prevention and rehabilitation.” Granted 18 June 2024; describes sensor-to-baseline comparison, rehabilitation-progress computation, alerts, and data-driven regimen suggestions. Patent document. Google Patents displayed George Mason Research Foundation, Inc. as assignee and “Active” as status when checked; neither entry is a legal opinion.