<?xml version="1.0" encoding="utf-8"?><feed xmlns="http://www.w3.org/2005/Atom" ><generator uri="https://jekyllrb.com/" version="3.10.0">Jekyll</generator><link href="https://josephruff.github.io//feed.xml" rel="self" type="application/atom+xml" /><link href="https://josephruff.github.io//" rel="alternate" type="text/html" /><updated>2026-10-02T20:45:39+00:00</updated><id>https://josephruff.github.io//feed.xml</id><title type="html">Joseph Ruff</title><subtitle>Computer Scientist / Software Developer</subtitle><author><name>Joseph Ruff</name></author><entry><title type="html">Multimodal Biosignal Preprocessing Pipeline</title><link href="https://josephruff.github.io//projects/multimodal-preprocessing-pipeline/" rel="alternate" type="text/html" title="Multimodal Biosignal Preprocessing Pipeline" /><published>2026-04-12T00:00:00+00:00</published><updated>2026-04-12T00:00:00+00:00</updated><id>https://josephruff.github.io//projects/multimodal-preprocessing-pipeline</id><content type="html" xml:base="https://josephruff.github.io//projects/multimodal-preprocessing-pipeline/"><![CDATA[<p>
	 This project implements a fully reproducible pipeline for downloading, preprocessing, and validating five open-access multimodal time-series datasets — PAMAP2, WISDM, mHealth, EEGMMIDB, and PTB-XL — in preparation for downstream self-supervised learning workflows. 
</p>

<p>
	 The pipeline harmonises three wearable activity recognition datasets (HAR) to a common 20 Hz representation with a shared six-channel accelerometer/gyroscope schema and unified class label taxonomy, enabling a single model to train across datasets. EEG data from the PhysioNet EEG Motor Movement/Imagery Database is preprocessed using MNE, with event-aligned 4-second epochs extracted from motor imagery runs. 12-lead ECG data from PTB-XL is ingested via the PhysioNet AWS S3 mirror, bandpass filtered, and split into patient-safe train, validation, and test folds. 
</p>

<p>
	 All outputs are stored as float32 NumPy arrays in a consistent [N, C, T] format alongside structured metadata CSVs covering subject provenance, label mappings, sampling rates, channel schemas, and QC flags. A validation script checks array integrity, label distributions, subject-level leakage controls, and HAR harmonisation across datasets. 
</p>

<p>
	 The pipeline scored 3rd out of 17 submissions in a competitive technical assessment for a Research Assistant post at Imperial College London, with the preprocessing plan rated the most thoroughly reasoned of all submissions. 
</p>

<hr>

<p style="text-align:center">
	<a href="https://github.com/JosephRuff/multimodal-preprocessing-pipeline" class="button icon brands fa-github">Source Code</a>
</p>]]></content><author><name>Joseph Ruff</name></author><summary type="html"><![CDATA[This project implements a fully reproducible pipeline for downloading, preprocessing, and validating five open-access multimodal time-series datasets — PAMAP2, WISDM, mHealth, EEGMMIDB, and PTB-XL — in preparation for downstream self-supervised learning workflows.]]></summary></entry></feed>