Independent research — how systems are trusted again after something goes wrong.
If you received an email from me and wanted to know who was asking before you answered, this page is that answer. It is the same question I would ask.
Software engineer, based in the San Francisco Bay Area.
Six years of my career have gone into test and validation platforms — the automated environments that decide whether a large hardware-and-software system is fit to ship. Building the harnesses, the regression suites, and the evidence that engineers and their managers actually rely on when they say a thing is ready. More recently, evaluation systems for machine-learning workloads: measuring whether an automated system did what it was meant to do, and how you would know if it had not.
Computer science at Purdue. The through-line of everything I have worked on is one question: how do you know it works?
Nobody. No university, no laboratory, no vendor, no sponsor. Nobody is paying for this and nobody is funding it. It is my own project, on my own time, out of my own pocket.
I do have a full-time job as a software engineer. It is unrelated to this work, has no involvement in it, and has no claim on it. I am not doing this on anyone's behalf.
Neither, and I would rather say so plainly than dress it up.
It is not academic research. There is no institution behind it, no ethics board, and no paper in submission, and I am not going to imply otherwise.
It is also not a sales process. There is no product, no company, and nothing to buy. I am a practitioner trying to understand a problem properly before deciding whether anything ought to be built for it. If that ever changes, you will not find out through one of these conversations — I do not turn research contacts into a pipeline.
Recovery after an automated action goes wrong. When an automated process or an agent writes something wrong to a production system of record: how does anyone notice, how long does it take to establish what happened, what does the recovery actually restore, what does it not, and who signs off that it is complete.
Validation after a change. How a cluster or a node is judged fit to take work again after something changed underneath it — a maintenance window, a driver or image update, a repaired node coming back. What the tests catch, and what a user ends up reporting first.
Both are the same question in different clothes: what evidence do you need before you trust a system again.
Twenty-five minutes. That is the whole ask.
No preparation, no materials, nothing to read beforehand. It is a conversation about how things actually work where you are — not a survey, and not a demo.
If you would like to pick a time directly: cal.com/harshit-singh-kehtyb. If none of the times there suit you, email me two that do and I will book around them.
The synthesis, before it goes anywhere else. Anonymised, aggregated across everyone who takes part, and sent to participants first.
Where this honestly stands: nothing is published yet. This is early, I am still building the picture, and I am not going to promise you a date I cannot keep. If you would rather wait until there is something to read before deciding whether to spend the time, say so — that is a fair thing to ask of a stranger, and I would rather you did that than sat through a call you regretted.