Sitelet https://github.com/bass990
Skip to content
View bass990's full-sized avatar

Block or report bass990

Block user

Prevent this user from interacting with your repositories and sending you notifications. Learn more about blocking users.

You must be logged in to block users.

Content in all repositories owned by your account will be closed.
Maximum 250 characters. Please don’t include any personal information such as legal names or email addresses. Markdown is supported. This note will only be visible to you.
Report abuse

Contact GitHub support about this user’s behavior. Learn more about reporting abuse.

Report abuse
bass990/README.md

Mamadou Bassirou Diallo

MSBA + AI at UT Dallas, graduating May 2027. I build AI systems and the data platforms under them, end to end, and I run the evaluation that could kill my own favourite feature before I ship it. Several of the repos below changed their production default because of what that eval found.

Every project here is complete: tests, CI, a Docker image, a section on what broke while building it, and a section on what it cannot do yet.


AI engineering

Project What it is What actually happened
chainpilot Supply-chain disruption agent: detects a stock-out or supplier delay, runs ten tool calls, has specialists argue when confidence is low, drafts RFQs and a Slack alert, and stops at a human-approval gate above a trust threshold that moves with rated outcomes. 204-run eval. Plain direct synthesis scored 0.735 strict accuracy against 0.382 with always-on deliberation, and my conditional gate was the least stable branch (its trigger flipped on 12 of 34 scenarios). Direct synthesis is the default; the gate is a flag.
triageiq ER triage decision support: a symptom specialist flags red flags, a senior-nurse synthesizer assigns the ESI level, care area and action checklist. Nurse override with a required reason, print report, hash-chained audit log. 270-run eval. A single prompt scored 100% ESI accuracy and produced no report at all on 12 of 99 attempts, all on ambiguous cases; the lean pipeline scored 98.9% and never failed to produce one. That completeness row is why lean is the default. Zero critical misses on both.
clauseguard Contract-conflict agent: extracts clauses from two PDFs, finds and ranks every conflict, drafts a redline brief, and has a second model review each suggested resolution. Live tool-call feed over SSE, React UI, Docker. 450-run eval on Sonnet 5. The agentic loop beat the single prompt by 5.4 F1 points, outside the run-to-run noise, so it became the default. Inlining the negotiation playbook into the prompt cost 14 points; exposing it as a tool was neutral. The router I had built did not earn its place and is now an opt-in cost lever.

Data engineering

Project What it is What actually happened
sec-filings-lakehouse Apache Iceberg lakehouse on real SEC EDGAR filings: 7.07M point-in-time financial facts (every restated version kept, "what did we know on date D" is one predicate), dbt with an enforced contract, 10-K/10-Q text chunked and embedded into pgvector as immutable index versions behind a retrieval-eval gate, FastAPI RAG with citations, Dagster. Docker Compose locally; Terraform applied and the pipeline verified on AWS, Azure and GCP. The gate rejected two of my three embedding indexes and was right both times: one was a retrieval problem (fixed with hybrid search), one was two bugs of mine in the chunker and the eval. 17,977 restated facts in two quarters is why the point-in-time model exists.
gharchive-streaming-lakehouse Kafka (Redpanda) with Avro and a schema registry, Spark Structured Streaming writing exactly-once into Iceberg, Debezium CDC folded into an SCD2 dimension, replay and backfill, Prometheus + Grafana, and a FinOps model that prices three AWS designs. Verified on AWS, Azure and GCP. 169,753 events landed with 169,753 distinct ids after a full-hour replay wrote zero duplicates. The first run put every event in 1970; a freshness gauge caught it, consumer lag did not. The README lists the five things that broke.
stackoverflow-causal-retention Causal inference on 1.77M Stack Overflow users: does a fast first answer make new contributors stay? BigQuery pipeline, 44 GB scanned for $0.22. Four estimators agree on about +7.7 points of 30-day retention. The instrumental-variable estimate came out at -20 points, which says the exclusion restriction does not hold, and the README reports it as a failure instead of dropping it.

Machine learning

Project What it is What actually happened
sba-loan-default-prediction XGBoost default-risk model on 890K SBA loans, AUCPR 0.567 against an 18% base rate, with a FastAPI service, PSI drift monitoring, golden-prediction tests, and a scoring app on Hugging Face. The ensemble weight search picked one model (1.0, 0.0) and the metadata says so. The cost-weighted threshold analysis found 0.35 lowers expected loss 22% versus the F1-optimal 0.66; the app now offers both. Calibration error is 0.21, and the app says that too.
NBA-Contract-Value-Analyzer LightGBM salary model with a leakage-safe time split, empirical 80% intervals, a staleness flag, drift monitoring, a static site, and a Streamlit dashboard on Hugging Face. The notebook's R² of 0.741 was early-stopping on the test season. The honest retrain is 0.733, and a determinism test keeps it there. Real salary data covers about 75 players because the salary sites sit behind Cloudflare; the synthetic fallback is disclosed on every page and in the API.

How I work

  • I write the evaluation before I trust the feature, and I keep the report that changed my mind in the repo.
  • Every README has a "what went wrong" section. The bugs were real, the fixes are in the code, and I would rather you read them there than find them in an interview.
  • Local first, then cloud: each system runs on a laptop with Docker Compose, and the cloud modules are applied and verified, not just written.

Stack

Python · SQL · Apache Iceberg · Spark Structured Streaming · Kafka / Redpanda · Debezium · dbt · Dagster · DuckDB · Terraform (AWS, Azure, GCP) · Prometheus / Grafana · FastAPI · React · Anthropic API · pgvector · XGBoost / LightGBM · scikit-learn · BigQuery · Docker · GitHub Actions

Education

  • M.S. Business Analytics and AI, The University of Texas at Dallas, May 2027
  • B.S. double major in AI Engineering and Management Information Systems, Sahmyook University, Seoul. Taught in Korean.

Languages

English · French (native) · Korean · basic Spanish

Contact

bassiroudiallo1305@gmail.com · LinkedIn

Looking for full-time AI engineering roles starting summer 2027, new-grad or early-career, with data engineering as a close second. Based in Dallas-Fort Worth, open to relocation.

Pinned Loading

  1. chainpilot chainpilot Public

    Supply-chain disruption agent with a human-approval gate. A 204-run eval chose direct synthesis over deliberation; the report is in the repo.

    Python

  2. triageiq triageiq Public

    ER triage decision support, lean pipeline chosen by a 270-run eval. Zero critical misses, every run produces a report. Mock patients only.

    Python

  3. clauseguard clauseguard Public

    Contract-conflict agent with a second-model review of every resolution. 450-run eval: the agentic loop beat the single prompt by 5.4 F1.

    Python 1

  4. sec-filings-lakehouse sec-filings-lakehouse Public

    Apache Iceberg lakehouse on real SEC EDGAR filings: 7.07M point-in-time facts, dbt contract, pgvector RAG index behind a retrieval-eval gate. Verified on AWS, Azure and GCP.

    Python

  5. gharchive-streaming-lakehouse gharchive-streaming-lakehouse Public

    Kafka (Redpanda) + Avro to Apache Iceberg with Spark Structured Streaming, exactly-once, Debezium CDC to SCD2, Prometheus/Grafana, FinOps. Verified on AWS, Azure and GCP.

    Python

  6. sba-loan-default-prediction sba-loan-default-prediction Public

    XGBoost default risk on 890K SBA loans, AUCPR 0.567, cost-weighted threshold and calibration disclosed, live scoring app.

    Jupyter Notebook