top of page
Abstract digital-interface-showing-real-time-data-streams-and-system-health-metrics-dashbo

Blind test: prepare your data

Same method, three kinds of loop. Pick yours, send two files, and score our result against your own records.

Robotic Arm Assembly

Robots

1. Two files

  • Normal baseline: recordings of the robot doing its usual cycles, no known faults. These calibrate the reference.

  • Test log: recordings that include faults or deviations you know about, labels removed. You keep the labels.

 

2. How much baseline

  • Ideal: about 1,000 normal cycles. For a 20-second cycle that is under 6 hours of normal operation, often a single shift.

  • Fewer is fine to start. Send what you have and we tell you after a first look whether it is enough.

 

3. What we need

  • Time and cycle

  • Joint position (measured and commanded)

  • Motor current

  • Torque

Example: one row per control sample, with timestamp, cycle, joint 1–6 position, current and torque.

 

4. Practical details

  • Sampling: 100 Hz or faster, as in our published result.

  • Format: CSV, Parquet or rosbag. Your own column names are fine with a short mapping note.

  • Length: each recording at least about 3 seconds, the earliest point a reading exists.

  • Anonymizing: rename signals and strip robot or site identifiers as you like.

 

5. What you get back

For each test recording: the detected timestamps and a severity, before you reveal your labels. You score the result against your own records.

Artificial Intelligence Circuit

RL agents

1. Two files

  • Normal baseline: episodes or runs where the agent behaved as intended.

  • Test log: episodes with degradations you know about (changed environment, sensor faults, a modified policy, disturbances), labels removed.

 

2. How much baseline

  • As many normal episodes as you can share.

  • Fewer is fine to start; we tell you after a first look whether it is enough.

  • Reward is not needed. The method does not read it.

 

3. What we need

  • Episode and step

  • State (what the agent observes)

  • Action (what the agent does)

Example: one row per step, with episode, step, state values and action values.

 

4. Practical details

  • Observations and actions as separate columns, not merged into one vector.

  • Actions that vary with the situation; a fixed replay or a constant command cannot be read.

  • A deployed or frozen policy; logs from while it was still training are not suitable.

  • Format: CSV, Parquet or JSONL. Your own column names are fine with a short mapping note.

  • Anonymizing: rename or rescale signals; we do not need to know what they mean.

 

5. What you get back

For each test episode: where the agent lost coupling with its environment (step or time) and a severity, before you reveal your labels.

Published result: on 21 frozen agents, the monitor caught 69% of perturbations versus 44% for reward-based monitoring.

Black Big Data_2.png
HDT_Human Perspective.jpg

LLM pipelines

1. Two files

  • Normal baseline: conversations or pipeline runs that went as intended.

  • Test log: conversations you know went off track (topic drift, loops, ignored instructions, a degraded model or tool), labels removed.

 

2. How much baseline

As many normal conversations as you can share. Fewer is fine to start; we tell you after a first look whether it is enough.

 

3. What we need

  • Conversation and turn

  • Who spoke (user, model, tool)

  • Message text

Example: one record per turn, with conversation, turn, role and text.

 

4. Practical details

  • Format: JSONL preferred, CSV fine.

  • Privacy: pseudonymize names, IDs and sensitive content; consistent replacement is enough.

  • Scope: no model weights, logits or API access, only the text that went in and out.

 

5. What you get back

For each test conversation: the turn where drift sets in and a severity, before you reveal your labels.

Published result: when each response is swapped for one from another turn, 87–92% of the measured link disappears while the surrounding prompts stay untouched.

Stock Market

Request a blind test

Tell us which section fits and roughly what data you have. We reply with how to transfer the files.

 

Email: idt@semarx.com

bottom of page