Signals researchers can trust.
Mohamed A M Elansary, PhD — scientific evaluation under uncertainty, production agentic LLM systems, and clear accounts of failure modes and evidence limits.
Scientific evaluation
- Designed multimodel forecast comparisons across basins and hydroclimates.
- Quantified uncertainty and validated imperfect USGS, NOAA, and NASA observations.
- Ran reproducible Python, R, Bash, Linux, and HPC workflows.
Production systems
- Builds GPT, Claude, and Gemini agent workflows at Vertexium.
- Uses regression evaluation sets for production agent behavior.
- Ships retrieval, routing, tenant isolation, provenance, and validation systems.
Evaluation approach
Define intended behavior and a failure taxonomy; label ambiguity and provenance; compare simple baselines; inspect disagreement and false signals; stratify results by user, task, and environment; then state precisely what the evidence does and does not support.
Honest fit boundary
I have not built LLM judges or led formal human-evaluation programs. I do not claim RLHF, benchmark-auditing, AI-safety research, or distributed large-model-training experience. My contribution is rigorous forecast evaluation, uncertainty quantification, scientific computing, and production agent regression evaluation.
Role and location
Research, Post-Training Evals · “This role is based in San Francisco, California.” Relocation with a support package is an honest discussion point.
Compensation is unpublished for this role. Official role posting