Thinking Machines Lab · Research profile

Signals researchers can trust.

Mohamed A M Elansary, PhD — scientific evaluation under uncertainty, production agentic LLM systems, and clear accounts of failure modes and evidence limits.

Model evaluationUncertainty quantificationScientific data + HPCAgent systems

Scientific evaluation

  • Designed multimodel forecast comparisons across basins and hydroclimates.
  • Quantified uncertainty and validated imperfect USGS, NOAA, and NASA observations.
  • Ran reproducible Python, R, Bash, Linux, and HPC workflows.

Production systems

  • Builds GPT, Claude, and Gemini agent workflows at Vertexium.
  • Uses regression evaluation sets for production agent behavior.
  • Ships retrieval, routing, tenant isolation, provenance, and validation systems.

Evaluation approach

Define intended behavior and a failure taxonomy; label ambiguity and provenance; compare simple baselines; inspect disagreement and false signals; stratify results by user, task, and environment; then state precisely what the evidence does and does not support.

Honest fit boundary

I have not built LLM judges or led formal human-evaluation programs. I do not claim RLHF, benchmark-auditing, AI-safety research, or distributed large-model-training experience. My contribution is rigorous forecast evaluation, uncertainty quantification, scientific computing, and production agent regression evaluation.

Role and location

Research, Post-Training Evals · “This role is based in San Francisco, California.” Relocation with a support package is an honest discussion point.

Compensation is unpublished for this role. Official role posting