Research notes · a camera-based screener for two Beighton hypermobility tests
Generalized joint hypermobility is usually screened with the Beighton score, a nine-point clinical exam a clinician performs by hand. Two of those nine points come from the hand alone: whether the little finger can be passively extended past 90 degrees, and whether the thumb can be bent down to touch the forearm. romwatch measures both of them through a webcam, entirely on-device, and says plainly what it can't do.
A Beighton score at or above a cutoff, combined with other criteria, is part of how hypermobile Ehlers-Danlos syndrome and generalized hypermobility spectrum disorder get diagnosed (Malfait et al., 2017). Most people never get screened for this at all: it isn't a routine part of a checkup, and the condition is often missed for years, particularly in people who read as merely "flexible" rather than as having a joint pattern worth mentioning to a doctor.
romwatch tries to close a small piece of that gap. A webcam watching a hand can measure both of the Beighton score's upper-limb items without needing to track a joint the camera can't see. It cannot replace the other seven points of a real exam (elbow and knee hyperextension, trunk flexion), and says so throughout, including here.
| Source | Year | What it establishes |
|---|---|---|
| Beighton, Solomon & Soskolne, Annals of the Rheumatic Diseases | 1973 | The original nine-point scoring system. Both of romwatch's maneuvers are drawn directly from it. |
| Grahame, Bird & Child (Brighton criteria), J. Rheumatology | 2000 | Ties the raw Beighton number to an actual diagnostic framework. |
| Malfait et al., Am. J. Med. Genet. C | 2017 | Current hEDS / G-HSD diagnostic criteria; requires an age-adjusted Beighton cutoff as one of several factors, not a diagnosis alone. |
| Malek, Reinhold & Pearce, Rheumatology International | 2021 | The score is upper-limb-weighted and shouldn't be used alone to rule generalized hypermobility in or out. Anchors most of the limitations below. |
| Juul-Kristensen, Rombaut et al., Am. J. Med. Genet. C | 2017 | Even clinician-administered Beighton scoring has only fair-to-poor reliability, a ceiling any camera-based version inherits. |
The state machine every attempt runs through, per maneuver.
Before any maneuver, the user holds a hand up naturally in front of the camera. Once MediaPipe's Hand Landmarker detects a stable hand for about half a second, the app records the diagonal of the box around all 21 landmarks, in pixels, as a personal reference size, used to judge camera distance for the rest of the session. An earlier version recorded bounding-box width instead, which measures how the hand is turned as much as how big it is: turning a hand edge-on collapses its width toward nothing while the hand stays exactly where it was. Since the thumb maneuver requires turning the hand, the reference stopped matching partway through every correct attempt. The diagonal barely moves under rotation.
Camera position is checked continuously, but only to catch cases where the measurement itself would be unreliable: no hand visible, part of the hand outside the frame (landmarks placed off-screen are extrapolations, not observations), or a hand so small or so large in frame that landmark placement degrades. It is explicitly not a pose check. Angle is free, because the measurements below are built not to care about it, and position in frame is free as long as the whole hand is inside it: an earlier version asked for the wrist to stay 10% clear of every edge, which flagged the natural habit of resting the wrist low in frame while missing fingers actually clipped off the top. During capture, frames that fail framing are excluded from scoring rather than merely counted against the attempt's quality, since each repetition keeps its most extreme frame and a clipped hand can produce an extreme-looking reading out of landmarks the model was guessing at.
While a maneuver is being performed, the number being scored is on screen in plain language, alongside the threshold it has to clear ("thumb tip sits 38% of a palm length out from the wrist; the line is at 10%"), and for the thumb maneuver the decision boundary itself is drawn on the camera overlay as a dashed line across the wrist, generated by the same function that scores the rep. The recorded value appears again next to each result. Any other hand in view is drawn faintly alongside it, so the choice of which hand is being measured is visible rather than assumed. This is a debuggability feature before it's a UI feature: the failure this app is most prone to is disagreeing with someone about their own body, and "it said no" is not a report anyone can act on, while "it said 38% while my thumb was flat against my wrist, and the dashed line was up near my knuckles" localizes the fault immediately.
Not every frame containing a hand can be scored, and the cases where it can't are specific enough to name. Calibration won't lock a reference while two hands are in view: that's the moment the session decides which hand it measures, the model's ordering between two detections is arbitrary, and picking the assisting hand there would mean measuring the wrong hand for the rest of the session with nothing on screen looking wrong. During a maneuver, a frame is dropped if the tracked hand isn't found, isn't fully in frame, is too small or large, or fails a check belonging to that specific maneuver.
A repetition tolerates 35% of its frames being dropped before it's discarded and retried. That budget was 15% when a second hand in frame wasn't expected. It has to be generous now for the same reason the tracking rules do: the hand doing the pushing spends much of the attempt on top of the joint being measured, so losing frames is the normal case rather than evidence of a bad attempt.
An earlier version overlaid an animated ghost skeleton on the live camera feed as a moving target to match. In practice it did the opposite of its job: a second hand-shaped outline superimposed on the user's own tracked hand read as a confusing second hand, not a guide, and drew attention away from the text instructions sitting right above it. It was removed. In its place, a real reference photo of the completed maneuver (cropped from a public-domain Wikimedia Commons image, see the app's Credits) is shown next to instructions that stay on screen for the whole maneuver instead of being overwritten by live tracking feedback.
Both maneuvers romwatch measures are passive tests. In the clinic the examiner pushes the joint into position; the question the test asks is what range the joint has when something else moves it, not what the person can reach unaided. Someone with hEDS reaching thumb to forearm with their other hand pushing isn't compensating for a failure to do it alone, that's the test being administered correctly. romwatch substitutes the person's own other hand for the clinician's, which makes two hands in frame the normal operating condition of the app rather than an edge case to tolerate.
An early version tracked only one hand, so the assisting hand could hijack tracking outright or flicker between the two frame to frame. Tracking then moved to up to two hands, following whichever wrist stayed closest to where the tracked hand was a moment ago. That held while both hands were visible and failed at the one moment that decides the result: at full apposition the assisting hand covers the hand it's pushing, the model frequently reports only the assisting hand, and "closest to where we were looking" then resolves to the wrong hand entirely, scoring the helper's pose as the person's range of motion. Tracking is now keyed to the handedness label of the hand calibrated at the start of the session. Proximity only breaks ties between two detections of the same handedness, and if the tracked hand isn't among the detections at all, the frame is reported as untracked rather than measured on whatever else is in view. The cost is a retry; the alternative was a confident wrong number.
For the thumb-to-forearm maneuver, the thumb tip is projected onto the hand's own axis (wrist to middle knuckle) and normalized by hand size, giving a signed reach: positive while the tip is still out over the palm, zero at the wrist line, negative once it has crossed onto the forearm. Two earlier versions measured plain distance instead, first to an approximated forearm ray and then straight to the wrist landmark, and both scored genuinely hypermobile hands negative. The ray version depended on the hand's 2D direction in frame correctly indicating where the forearm was, and broke down for exactly the reason a real attempt tends to break it: a full thumb-to-forearm touch rotates the wrist away from wherever it started, taking the assumed ray with it. Distance to the wrist landmark removed that assumption but measured the wrong quantity, because the thumb reaches the forearm on the radial side and a few centimetres proximal to the landmark: a real touch still reads as a sizeable distance, while a threshold loose enough to accept it also accepts an ordinary thumb folded across the palm. The two cases barely differ in distance to the wrist and differ completely in whether the tip crossed the wrist line, which is what the clinical criterion asks about.
For the little-finger maneuver, the angle between the wrist-to-pinky-knuckle vector and the pinky's proximal phalanx is measured: how far the finger has folded back relative to the hand's own orientation, rather than relative to the camera frame. Both measurements are ratios between tracked landmarks, so neither depends on an assumption about which way the hand or arm is turned toward the camera. Both are computed after rescaling the landmarks to undo the frame's aspect ratio, since MediaPipe normalizes x by frame width and y by frame height: on a 16:9 frame one normalized unit means different real distances on the two axes, which skews any direction-sensitive measurement by however the hand happens to be turned. Undoing it first is what lets the angle-independence claim survive contact with a real webcam.
Each maneuver gets one attempt, and a second only if the first wasn't a clean positive. A repetition is discarded and retried (up to twice) if tracking quality drops during capture, rather than being scored as a miss. A single positive repetition is enough to score the maneuver positive for that session, matching how the real exam scores it: the clinician does the maneuver once per side, not several times looking for a majority. Session results are stored in the browser's local storage, and the on-screen recommendation escalates only once a maneuver has come back positive across three or more separate sessions, not on a single reading.
The version described above is the result of one round of testing by a person with a clinical hEDS diagnosis, not a study, a single test session's worth of honest feedback. It's included here because it's the reason several design decisions above exist, and because a portfolio piece that only shows the polished current state hides how it got there.
The first working version scored that same test session as within the typical range on both maneuvers, and the version after it still did. Four concrete problems came out of digging into why:
| Problem | Root cause | Fix |
|---|---|---|
| Tracking broke when the assisting hand entered frame | Beighton maneuvers are normally administered with a second hand pushing the joint into position. With tracking limited to one hand, whichever hand the model locked onto could flip between the test hand and the assisting hand frame to frame. | fixed Track up to two hands; follow whichever one stays closest to the last known position. |
| The thumb measurement assumed a hand direction a real attempt rotates away from | The forearm-ray approximation depended on the hand's 2D direction in frame, which a genuine full thumb-to-forearm motion changes as the wrist turns. | fixed Measure along the hand's own axis instead, which doesn't depend on hand direction. |
| A thumb visibly touching the wrist still scored negative | Distance from the thumb tip to the wrist landmark stays large even at full contact, because contact happens off to the radial side and past the landmark, so no threshold on that distance separates a hypermobile thumb from an ordinary one folded across the palm. | fixed Score signed reach along the hand's axis: did the tip cross the wrist line, not how far it sits from a point. The line is now drawn on the overlay and the live number shown on screen. |
| Three repetitions asked too much | Three attempts per maneuver, needing two positive, was tiring to perform and no more accurate than accepting one clean demonstration, which is what the clinical exam itself does. | fixed One attempt, a second only if the first wasn't a clean positive. |
What this single session doesn't establish: whether the fixed version reads correctly across different people, hand sizes, skin tones, lighting conditions, or camera setups. One person finding and describing four specific failure modes is enough to fix those four failure modes; it isn't evidence the tool is now accurate in general, and the limitations below still apply in full.
As a portfolio project, romwatch demonstrates a full pipeline: real-time hand tracking in the browser, geometry derived from real clinical tests rather than an invented metric, a capture and quality-gating state machine, and a consistency layer across sessions, all running client-side with no backend. As a clinical tool, it's exactly what its own disclaimer says: a screening aid inspired by two items of a real clinical score, not a diagnostic device. Turning it into something closer to clinically useful would mean, at minimum, validating its thresholds against a clinician-scored cohort and extending it toward the other seven Beighton items that a single hand-tracking camera can't reach on its own.