← The Lab

03 · Security × ML · 2025

Credential stuffing detection

Graph auto-encoders over real production auth logs — users, IPs and user-agents as a heterogeneous graph, anomalies scored by reconstruction error. Test AUC 0.93 / AP 0.96.

Scatter plot of learned graph embeddings: grey user clusters, IP nodes coloured from light to dark red by anomaly score, with the most anomalous IPs grouped in a separate cluster on the right.
t-SNE of the embeddings learned by the improved graph auto-encoder: users in grey, IPs coloured by anomaly score. Figure from the project repository.

Engineering project at Télécom Saint-Étienne, built on real, pseudonymized production authentication logs from an industry partner. The goal: detect credential stuffing, distributed brute force and account-takeover patterns.

Rule-based detection reads events one at a time, but these attacks are relational — one IP against many accounts, one account hit from many IPs and user-agents. Login events are modelled as a heterogeneous graph of users, IPs, user-agents and countries, and graph auto-encoders (GAE and VGAE) score anomalies by how badly they reconstruct each link.

Structure alone was not enough: the original model reached a test AUC of 0.37. Injecting behavioural features into the nodes — connection frequency, user-agent diversity, per-IP failure rate — raised it to 0.93, with an average precision of 0.96. The data stays private; the repository holds the full method.