Separating Knowledge from Psychomotor Skill
Knowing a procedure and being able to perform it are different constructs. A trainee may master the anatomy, energy modalities and complication pathways of laparoscopic surgery on paper, yet lack the manual dexterity to operate safely. Well-designed curricula therefore assess theory and psychomotor ability separately, recognising that a written examination cannot predict how steadily a candidate manipulates instruments inside the abdomen.
Case logbooks record exposure, not competence. Counting procedures says little about how safely each was performed, because supervised volume varies with rota, case mix and local practice. Objective, station-based testing addresses this by producing criterion-referenced metrics — task time, precision and error counts — that are reproducible between candidates and largely independent of an individual examiner's subjective impression on the day.
The GESEA programme, developed by the European Society for Gynaecological Endoscopy, formalises this separation. A theoretical test evaluates knowledge, while three psychomotor stations — LASTT for coordination and navigation, SUTT for suturing and knot-tying, and HYSTT for hysteroscopy — assess technical execution on validated bench models across levels of increasing difficulty.
This structure reflects a wider shift toward proficiency-based progression, in which a candidate advances only after demonstrating a defined standard rather than after a fixed period of attendance. A certificate under such a framework attests that structured assessment was completed to criterion; it does not, by itself, confer a licence to operate independently or certify unsupervised competence.
SUTT: Suturing and Intracorporeal Knot-Tying
SUTT — the Suturing and Knot-Tying Test and Training station — assesses one of the most demanding skills in minimal-access surgery. Intracorporeal suturing combines precise needle handling, accurate tissue apposition and the formation of secure knots, all performed through fixed ports with long instruments. It draws on the coordination measured by LASTT but adds a further layer of fine motor control.
The task is usually performed on a bench model, with the candidate placing stitches and tying knots to a defined standard within a time limit. Scoring considers completion time, knot security, suture placement accuracy and tissue handling. Because these dimensions can be measured objectively, the station again separates capability from confidence, distinguishing a surgeon who can reliably close tissue from one who merely reports having done so.
Suturing is deliberately assessed at a higher level than basic coordination because its failure modes carry real clinical weight: a slipped knot or an ischaemic bite has direct consequences for haemostasis and healing. Practising these movements on a model, to criterion, before applying them to patients aligns with the wider surgical-safety principle of building competence away from the operating table.
Structured suturing assessment also creates a shared vocabulary between trainer and trainee. When feedback references a specific, measured deficit — needle-angle control, for instance, rather than a general sense that a knot 'looked loose' — remediation becomes targeted and the trainee can rehearse the exact component that limits performance, then re-test against the same objective standard.
HYSTT: Hysteroscopy Skills
HYSTT — the Hysteroscopy Skills Testing and Training station — addresses the distinct demands of operating within the uterine cavity. Unlike laparoscopy, hysteroscopy works in a confined, distended space that depends on continuous fluid management, and the surgeon navigates and treats through a single channel. The perceptual and motor skills therefore differ enough to warrant assessment in their own right.
On a validated model, candidates practise orientation within the cavity, systematic inspection of the walls and ostia, and targeted manipulation such as reaching defined points or removing simulated lesions. Assessment records accuracy, economy of movement and the time taken to complete each task, again yielding objective metrics that separate a steady, systematic operator from one who searches the cavity haphazardly.
Because hysteroscopic complications — fluid overload, perforation or incomplete treatment — often arise from poor orientation and rushed technique, rehearsing these skills on a bench before the clinical setting is a reasonable safeguard. Including hysteroscopy alongside laparoscopy also reflects the breadth of endoscopic gynaecology, ensuring that a certification pathway does not measure only the abdominal component of the specialty.
Construct Validity, Simulators and the Operating Room
An assessment is only useful if it measures what it claims to. Construct validity — the capacity of a test to distinguish groups that genuinely differ, typically novices from experts — is the property that lets station scores stand as evidence of skill. Validated laparoscopic and hysteroscopic stations have repeatedly shown this discrimination, which is why they can reasonably gate progression through a structured pathway.
Both physical box trainers and virtual-reality simulators support this model. Box trainers offer authentic instruments, real tissue feel and low cost; VR systems add automated metrics and standardised scenarios without consumables. Neither is inherently superior — the evidence supports both as valid environments for acquiring and measuring skill — and many programmes use them in combination according to the competency being assessed.
The essential contrast is with unstructured operating-room exposure, where learning is opportunistic and uneven. Theatre teaching remains indispensable, yet as an assessment environment it is confounded by case difficulty, patient factors and supervisory style — variables no examiner fully controls. Objective station testing, by design, holds these constant, and differs from ad-hoc theatre learning in several important respects:
- Standardisation — every candidate faces the same task under the same conditions, so scores are comparable.
- Reproducibility — exercises can be repeated to chart a learning curve rather than captured once by chance.
- Objective metrics — time, accuracy and error are recorded directly, reducing reliance on subjective recall.
- Patient safety — skills are acquired and failures absorbed on models, not on patients.
- Focused feedback — a measured deficit points to the specific component that needs rehearsal.
