Open to Research & Applied Scientist roles

Katarzyna
Szymaniak

Machine learning for signals that won't hold still.

Choose your field — the page adapts
Tailored for wearables & health. Switch any time.

I build biosignal and wearable models — sEMG, IMU, EEG — that still work on the next user, the next session and the next week. Not just on the benchmark. I build models for noisy, non-stationary, multichannel time series that hold up when the data drifts and labels are scarce — and the evaluations that prove it. New speakers, new microphones, new rooms: audio has the same shift problem I solved on biosignals. I'm now bringing those methods to audio and audio-language models.

PhD · Edinburgh · 2026Ex-Meta Reality Labs (CTRL-labs)Collaborating with NVIDIA
01 — About

One problem, many signals

Most ML works when the world holds still. Real signals don't — so I build the models that keep up, and the evaluations that prove it.

I'm an ML researcher with a PhD from the University of Edinburgh (January 2026), supervised by Prof. Kianoush Nazarpour and Prof. Timothy Hospedales. My thesis — Mitigating Confounding Factors in Myoelectric Control Through Adaptive Modelling and Learning — sits where deep learning meets the inconvenient reality of biosignals: electrodes shift, postures change, and users adapt to the model while the model is trying to adapt to them.

At Meta Reality Labs (CTRL-labs) in New York, I took that problem to wrist-worn neural interfaces: fusing sEMG with IMU, and building few-shot personalisation for a Conformer decoder that had to hold up when the band came off and went back on.

My research targets the failure modes that decide whether a model survives deployment on continuous sensor data: non-stationarity, cross-user and cross-session shift, labels that cost human time, and signals corrupted by noise and motion. I've attacked them from three sides — a domain-generalisation model that needs zero data from the unseen condition, active learning that replaced a full calibration session with 3.2 minutes of queried labels, and benchmarks rigorous enough (~40k models, nested cross-validation) to separate real gains from noise.

I work end-to-end, from designing the human-subject protocol and collecting the data to modelling and statistical evaluation. None of this is specific to EMG: speaker variability, device mismatch and low-resource adaptation in audio share the same structure, and that is where I'm taking the work next.

I'm looking for the team to do this with in earnest — biosignal or audio foundation models, health AI, neural interfaces, or any group that treats real-world robustness as the research problem, not the last ablation.

Same problem, different signal

The failure modes I work on, as they appear in each domain — and the evidence I have for each. The highlighted column follows the lens you picked above.

Problem
Wearables
Time series
Audio
What I've shown
The world moves under the model
Electrode shift, arm posture, sweat, fatigue
Regime change, sensor drift, broken seasonality
New room, noise floor, microphone
Quantified a ≤10% position-induced drop (GREAT); generalised to unseen positions with zero target data (DR-cGMM).
New user, new device
Cross-user, cross-session sEMG; re-donning
New sites, machines, patients
New speakers, accents, hardware
Few-shot per-user calibration of a Conformer decoder at Meta; cross-session learning curves over calibration budgets.
Labels are the bottleneck
User calibration time, clinic time
Expert annotation budgets
Transcription and preference labels
3.2 min of actively queried data beat a full calibration session (86% vs 82%).
More than one stream
sEMG + IMU at different rates
Multivariate sensors, asynchronous sampling
Audio + text
Early-fusion sEMG + IMU decoders; diagnosed modality imbalance — one stream dominating optimisation.
Benchmarks flatter the model
Within-session splits
Look-ahead and cross-domain leakage
Test-speaker and test-device leakage
Leave-one-domain-out model selection, nested CV and Friedman/Nemenyi testing over ~40k models.
02 — Research

Three questions

Everything I work on comes back to three questions about deployed models on continuous signals. Each has peer-reviewed evidence behind it — and each transfers beyond the signal I first answered it on.

01 · Robustness

Distribution shift

How do you build a model that works in a condition it has never seen?

I treat recording conditions as domains and design for the unseen one: a class-conditional Gaussian mixture with a covariance regulariser shared across domains, which needs zero data from the target condition.

In your domain Arm position, band re-donning, long-term signal drift. Regime changes, new sites, slowly drifting sensors. Unseen microphones, rooms and noise conditions.
  • DR-cGMM — 94.3% vs 93.4% ERM; top-ranked for 6 of 8 users
  • GREAT — an open benchmark that isolates one controlled shift
02 · Label efficiency

Adaptation on a budget

What is the smallest labelling budget that still moves the model?

Acquisition functions as a research object, not a utility: uncertainty-based querying (least-confidence, margin, entropy), few-shot personalisation, and parameter-efficient fine-tuning when labels cost human time.

In your domain Minutes of user calibration instead of full sessions. Spending expert annotation where it changes the model. Selecting SFT and DPO data when labels are scarce.
03 · Rigour

Evaluation you can trust

Would this result survive deployment?

Splits that mirror deployment, model selection that never peeks at the test domain, and non-parametric significance tests — so that "method X helps" means something.

In your domain Cross-subject and cross-session splits by design. No look-ahead, no leakage across domains or entities. Held-out speakers, devices and acoustic conditions.
  • ~40k trained models, leave-one-domain-out + nested CV
  • Showed popular deep DG methods give no significant gain over ERM in the low-data regime
03 — Now

Where I'm heading

New direction · 2026
Early-stage · in collaboration with NVIDIA

Robust adaptation of audio foundation models

Audio is the next test bed for the question behind all my work: what happens when a pretrained model meets a speaker, device or room it has never encountered? Audio-language models are powerful but brittle under acoustic shift, and adapting them usually assumes plenty of labels. I'm starting from three questions:

Adapt
When does lightweight fine-tuning (LoRA, SFT) match full adaptation with only a handful of labels?
Stress-test
Do the gains hold on held-out speakers, devices and acoustic conditions, measured without leakage?
Spend labels wisely
Can uncertainty-guided selection stretch SFT and preference data, as active learning did for EMG calibration?
Post-training Audio-language models Shift-aware evaluation

Questions on my list

Do audio LLMs actually listen?

Reasoning-tuned audio models may lean on text rather than sound as their chains of thought grow. Where does attention to the audio fade, and can training-free re-grounding bring it back?

When is reasoning worth the compute?

A router that decides, per input, whether extra inference helps or hurts — treating compute as an acquisition problem rather than a fixed budget.

Same lens as the wearables work: distribution shift, budgets, and the evaluation that decides whether a model survives contact with deployment.

04 — Industry

Research at product scale

Meta Reality Labs · CTRL-labs · New York

Research Scientist Intern, Continuous Navigation

Sep 2022 – Mar 2023 · Host: Fabio Stefanini

The team's cross-user sEMG decoding research was published in Nature (2025) and underpins the Meta Neural Band.

  • Multimodal fusion. Extended the wrist-worn sEMG decoding stack to sEMG + IMU: inertial preprocessing, early-fusion decoders, and diagnosis of modality imbalance between the two streams.
  • Few-shot personalisation. Built an end-to-end pipeline for new custom gestures — from the human-subject data-collection protocol to few-shot per-user calibration of a Conformer-based decoder (the convolution-augmented Transformer from speech recognition).
  • Evaluation under shift. Measured cross-session generalisation after the band is re-worn, with per-user learning curves over calibration budgets; presented results and limitations at the team's technical deep dive.
05 — Papers

Selected publications

EurIPS 2025Workshop · extended version under review at TMLR

Position: Domain Generalisation in Myoelectric Control

K. Szymaniak, K. Nazarpour, H. Gouk

Proposed DR-cGMM, a domain-regularised class-conditional Gaussian mixture that generalises to unseen domains with zero data from them. It was the only method to significantly outperform ERM, MMD and CORAL; popular deep DG methods gave no significant gain over ERM. Leave-one-domain-out model selection with nested CV and Friedman/Nemenyi testing over ~40k models.

Domain generalisationBenchmarkingProbabilistic models
Scientific Data 2025Open dataset

EMG Dataset for Gesture Recognition with Arm Translation (GREAT)

I. Kyranou*, K. Szymaniak*, K. Nazarpour  * equal contribution

Designed the acquisition protocol for an open benchmark of 16-channel sEMG at 2 kHz with synchronised 18-sensor hand kinematics over a 3×3 grid of arm positions. Quantified the shift — up to 10% accuracy drop, and position itself decodable from the signal (~80%) — and showed position-aware hierarchical decoding does not fix it outside an i.i.d. setting.

Open dataControlled shiftsEMG
Front. Neurorobotics 2022

Recalibration of Myoelectric Control with Active Learning

K. Szymaniak, A. Krasoulis, K. Nazarpour

First application of active learning to myoelectric control. 3.2 minutes of actively queried data beat a full calibration session (86% vs 82%; 12 able-bodied participants, ~2 points in 2 transradial amputees; offline simulation). Identified a failure mode: diversity-based batch selection degraded margin sampling, likely because informative boundary samples cluster and get discarded as redundant.

Active learningHuman-in-the-loop
IEEE TNSRE 2026Survey

Deep Feature Learning from Electromyographic Signals for Gesture Recognition Systems

W. Zhong, X. Jiang, K. Szymaniak, M. Jabbari, C. Ma, K. Nazarpour

Survey of deep representation learning for EMG-based gesture recognition.

Earlier research
  • Predicting behaviour from brain data (Edinburgh, 2021; Prof. Oisin Mac Aodha, with the Centre for Discovery Brain Sciences) — decoded the location of freely behaving mice from in vivo deep-brain calcium imaging, with neuron-activation maps inspired by class activation mapping.
  • Sequence models for EEG in decision deadlocks (Swansea, 2020; Dr Jingjing Deng, Prof. Xianghua Xie) — temporal dependencies, non-stationarity and cross-channel connectivity for brain–computer interfaces.
  • HoloLens neuroimaging for medical education (Swansea, 2019; Best Student Project) — MRI brain scans as interactive 3D meshes in mixed reality. Demo ↗
06 — Background

The trajectory

2026 →

Independent research, audio-language models

In collaboration with NVIDIA
2021 – 2026

PhD, Biomedical Artificial Intelligence

University of Edinburgh, School of Informatics · Supervisors: Prof. Kianoush Nazarpour, Prof. Timothy Hospedales · Thesis: Mitigating Confounding Factors in Myoelectric Control Through Adaptive Modelling and Learning
2022 – 2023

Research Scientist Intern

Meta Reality Labs (CTRL-labs), New York
2022

Oxford Machine Learning Summer School

University of Oxford
2020 – 2021

MRes, Biomedical Artificial Intelligence

University of Edinburgh
2016 – 2021

MRes Visual Computing · BSc (Hons) Computer Science, First Class

Swansea University
07 — Writing

Notes & essays

08 — Contact

Let's talk

The next role.
A collaboration.
A good problem.

I'm looking for Research Scientist and Applied Scientist roles where models meet messy, real-world signals:

Wearables & health AI Neural interfaces Time-series & sensor ML Audio & speech adaptation Robustness & evaluation

Especially teams building biosignal foundation models, personalisation, or validated on-device ML — where cross-user and cross-session shift is the product problem. Especially teams whose data drifts — sensors, operations, clinical records — and who care whether a model still works next quarter. Especially teams adapting speech or audio-language models to new devices, domains and low-resource settings. Based in Edinburgh, open to relocation.