Adversarial ML
Attacks on model integrity rather than model accuracy: evasion that survives preprocessing, training-data poisoning, hidden backdoor triggers, and weight-level manipulation that leaves the loss curve intact.
est. independent security research
Eigenmeier is a security research firm. We break machine learning systems on purpose — their weights, their supply chain, their agent loops — then publish what we find and ship the tooling so you can reproduce it.
Representative output from our open-source artifact scanner.
01 Research
We work where model behaviour meets real infrastructure — because that is where the security boundary actually lives. Every track ends in a paper, a patch, or a tool.
Attacks on model integrity rather than model accuracy: evasion that survives preprocessing, training-data poisoning, hidden backdoor triggers, and weight-level manipulation that leaves the loss curve intact.
The layer most teams never audit. We audit serialized artifacts, tokenizer files, inference runtimes, container images, and plugin manifests for unsafe deserialization, remote fetch closures, and provenance gaps.
Tool-using systems have a privilege boundary that nobody drew. We study injection through tool output, memory and retrieval poisoning, credential leakage into context, and escalation across delegated sub-agents.
A benchmark you cannot reproduce is decoration. We build fixed-seed harnesses, pin prompts and decoding parameters, publish full transcripts, and measure whether a claimed mitigation actually moved the number.
02 Open source sample index
Research that ships with its reproducer. Apache-2.0 unless noted — fork it, audit it, break it.
Static analysis for ML model artifacts. Catches unsafe deserialization, obfuscated payloads, and provenance gaps before a checkpoint ever loads.
A policy firewall for LLM tool use. Constrains what an agent may call, validates arguments against a schema, and redacts secrets before they reach context.
Reproducible red-team benchmark runner. Fixed seeds, pinned system prompts, versioned transcripts — so a score means the same thing six months later.
Membership inference and extraction probes expressed as spectral analysis of a model's response surface, with confidence intervals rather than point estimates.
Gradient-guided fuzzer for robustness testing on vision and text encoders, with transfer rates measured against held-out surrogate architectures.
Checkpoint provenance end to end: build, sign, mirror, serve — with a verifiable chain and a diff view when a shard changes.
Traces agent loops and renders the instruction/tool-call graph, so injection paths through tool output become visible instead of theoretical.
Manifest linter for inference containers and model plugins. Flags writable mounts, network egress, and host paths before they reach a serving cluster.
No projects in this category yet.
03 Advisories sample index
We report to the maintainer first, publish when the fix ships or the deadline passes, and never trade a finding for silence.
04 About
We are an independent firm — no product to protect, no roadmap to satisfy. The only thing our reputation is built on is being right in public.
Findings, transcripts, and code leave the lab. If it is useful to an attacker it is useful to a defender first.
Ninety-day disclosure windows, vendor-friendly timelines, and credit where it is due. We do not embargo past a deadline.
A finding nobody can reproduce is a rumour. Every report lands with code, seeds, and the exact environment it ran in.
05 Contact
Briefing a security review, reporting a vulnerability, or just arguing about benchmarks — all three land in the same inbox and reach a human.