semarx


Robots
1. Two files
-
Normal baseline: recordings of the robot doing its usual cycles, no known faults. These calibrate the reference.
-
Test log: recordings that include faults or deviations you know about, labels removed. You keep the labels.
2. How much baseline
-
Ideal: about 1,000 normal cycles. For a 20-second cycle that is under 6 hours of normal operation, often a single shift.
-
Fewer is fine to start. Send what you have and we tell you after a first look whether it is enough.
3. What we need
-
Time and cycle
-
Joint position (measured and commanded)
-
Motor current
-
Torque
Example: one row per control sample, with timestamp, cycle, joint 1–6 position, current and torque.
4. Practical details
-
Sampling: 100 Hz or faster, as in our published result.
-
Format: CSV, Parquet or rosbag. Your own column names are fine with a short mapping note.
-
Length: each recording at least about 3 seconds, the earliest point a reading exists.
-
Anonymizing: rename signals and strip robot or site identifiers as you like.
5. What you get back
For each test recording: the detected timestamps and a severity, before you reveal your labels. You score the result against your own records.

RL agents
1. Two files
-
Normal baseline: episodes or runs where the agent behaved as intended.
-
Test log: episodes with degradations you know about (changed environment, sensor faults, a modified policy, disturbances), labels removed.
2. How much baseline
-
As many normal episodes as you can share.
-
Fewer is fine to start; we tell you after a first look whether it is enough.
-
Reward is not needed. The method does not read it.
3. What we need
-
Episode and step
-
State (what the agent observes)
-
Action (what the agent does)
Example: one row per step, with episode, step, state values and action values.
4. Practical details
-
Observations and actions as separate columns, not merged into one vector.
-
Actions that vary with the situation; a fixed replay or a constant command cannot be read.
-
A deployed or frozen policy; logs from while it was still training are not suitable.
-
Format: CSV, Parquet or JSONL. Your own column names are fine with a short mapping note.
-
Anonymizing: rename or rescale signals; we do not need to know what they mean.
5. What you get back
For each test episode: where the agent lost coupling with its environment (step or time) and a severity, before you reveal your labels.
Published result: on 21 frozen agents, the monitor caught 69% of perturbations versus 44% for reward-based monitoring.


LLM pipelines
1. Two files
-
Normal baseline: conversations or pipeline runs that went as intended.
-
Test log: conversations you know went off track (topic drift, loops, ignored instructions, a degraded model or tool), labels removed.
2. How much baseline
As many normal conversations as you can share. Fewer is fine to start; we tell you after a first look whether it is enough.
3. What we need
-
Conversation and turn
-
Who spoke (user, model, tool)
-
Message text
Example: one record per turn, with conversation, turn, role and text.
4. Practical details
-
Format: JSONL preferred, CSV fine.
-
Privacy: pseudonymize names, IDs and sensitive content; consistent replacement is enough.
-
Scope: no model weights, logits or API access, only the text that went in and out.
5. What you get back
For each test conversation: the turn where drift sets in and a severity, before you reveal your labels.
Published result: when each response is swapped for one from another turn, 87–92% of the measured link disappears while the surrounding prompts stay untouched.

Request a blind test
Tell us which section fits and roughly what data you have. We reply with how to transfer the files.
Email: idt@semarx.com