(ICML 2024) Pursuing Overall Welfare in Federated Learning through Sequential Decision Making (https://arxiv.org/abs/2402.14650)
-
Updated
Oct 9, 2024 - Python
(ICML 2024) Pursuing Overall Welfare in Federated Learning through Sequential Decision Making (https://arxiv.org/abs/2402.14650)
StarCraft II high-level feature extractor and replay visualizer
Machine-checked Lean 4 / Mathlib formalization of Ismail's Primitives — six structural primitives proven necessary, mutually independent, and sequentially linked for sequential decision-making under uncertainty. 0 sorry · 0 axiom · ~12,700 lines.
code for published paper titled "Optimal CO2 storage management considering safety constraints in multi-stakeholder multi-site GCS projects: a Markov game perspective "
A project to classification.
LLM benchmark for long-horizon sequential decision-making — language models play the turn-based strategy game LGeneral against the CPU over a REST API, scored on a graded victory scale.
Explainable action ranking from historical sequences — a Rust library and CLI, not a contextual bandit or online-learning library.
Research portfolio in Scientific AI: scientific machine learning, neural operators, digital twins, uncertainty quantification and sequential decision-making.
Reproducible research tools for result-aware decision workflows and conformal reference retention.
Implementation of a Decision Transformer for power-grid control in L2RPN. The model learns from offline Tutor/Junior demonstrations and continues training online, serving as a drop-in replacement for PPO in the original pipeline proposed by aspirin96.
Do humans adapt their planning horizon? - An analysis of sequential decision-making in the videogame Frogger
Four policy classes (PFA/CFA/VFA/DLA) on the energy-storage problem, scored against a verified exact optimum, with a model-misspecification study
Sequential finite-population certification on GB1 and AAV2 with complete paired-trial audit artifacts.
Autonomous Blue Team agent. Using Deep Reinforcement Learning for sequential intrusion detection and active incident response in simulated enterprise networks (CybORG / CAGE). Moving beyond static classification, this agent learns an interactive, explainable policy to continuously mitigate cyber threats.
砂場
A reproducible falsification harness for resource-bounded AI-agent scheduling under correlated failures, noisy verification, costs, deadlines, and false-acceptance constraints.
Researching policy-induced observability in recommender systems: when serving policies create self-confirming preference models, and when exploration is worth the cost.
Testing temporal reasoning in trading agents under random liquidation horizons
Mathematical study notes on Markov decision processes and dynamic programming, adapted from Puterman. Covers finite- and infinite-horizon MDPs, Bellman equations, policy evaluation, value iteration, policy improvement, and convergence proofs using Banach’s fixed-point theorem.
To associate your repository with the sequential-decision-making topic, visit your repo's landing page and select "manage topics."