All papers are listed below in reverse chronological order in which they appeared online.
Filters
All 2026 2025 2024 2023 2022 2021 2020 2019 2018 2017 2016 2015 2014 2013 2012 and earlier
All Acceleration Adaptive Asynchronous Broximal Compressed communication Coordinate descent Convex Decentralized Error feedback Federated learning First-order Linear algebra LLM Local training LoRA Minimax Momentum Muon Non-convex Primal-dual Privacy Proximal Second-order Sketching Stochastic Variance reduction Zero-order
Prepared in 2026
Understanding MARS: when scaling momentum provably helps
43rd International Conference on Machine Learning (ICML 2026)
Algorithms:
MARSGeneral analysis of LMO-based optimizers: beyond bounded variance
43rd International Conference on Machine Learning (ICML 2026)
Algorithms:
NSGD with momentum, MuonSuper-Tuning: from activation-aware pruning to sparse fine-tuning
arXiv github
Algorithms:
Super, SupraSpecGradFilter: a spectral gradient filtering framework for taming federated heterogeneity
arXiv
Algorithms:
SpecGradFilterConvergence analysis of Muon-type methods with inexact LMO in the degenerate case
arXiv
Algorithms:
inexact Gluon, inexact Gluon with weight decaySILAGE: memory-efficient, full-gradient-free nonconvex optimization for nested finite sums
arXiv
Algorithms:
SILAGEDemystifying pipeline parallelism: first theory for PipeDream
arXiv
Algorithms:
PipeDream, Randomized PipeDreamA unified primal-dual recipe for accelerating three-operator splitting methods
arXiv
Algorithms:
ACV-I, ACV-II, APDTR-I, APDTR-IILOSCAR-SGD: local SGD with communication-computation overlap and delay-corrected sparse model averaging
arXiv
Algorithms:
LOSCAR-SGDDistance-aware Muon: adaptive step scaling for normalized optimization
arXiv
Algorithms:
DA-Muon, SC-Muon, DF-MuonRingmaster LMO: asynchronous linear minimization oracle momentum method
arXiv
Algorithms:
Ringmaster LMORescaled asynchronous SGD: optimal distributed optimization under data and system heterogeneity
arXiv
Algorithms:
Rescaled ASGDRennala MVR: improved time complexity for parallel stochastic optimization via momentum-based variance reduction
arXiv
Algorithms:
Rennala MVRLocal LMO: constrained gradient optimization via a local linear minimization oracle
arXiv slides
Algorithms:
Local LMOBroximal alignment for global non-convex optimization
arXiv
Algorithms:
BPMCommunication-efficient Gluon in federated learning
arXiv
Algorithms:
Compressed Gluon with Error Feedback and MVRA Nesterov-accelerated primal-dual splitting algorithm for convex nonsmooth optimization
arXiv
Algorithms:
APAPCStabilized proximal point method via trust region control
arXiv
Algorithms:
TRPPMByzantine-robust and differentially private federated optimization under weaker assumptions
42nd Conference on Uncertainty in Artificial Intelligence (UAI 2026)
arXiv
Algorithms:
Byz-Clip21-SGD2MBiCoLoR: Communication-efficient optimization with bidirectional compression and local training
arXiv
Algorithms:
BiCoLoR
Prepared in 2025
First provable guarantees for practical private FL: beyond restrictive assumptions
arXiv
Algorithms:
Fed-α-NormEC, DP-Fed-α-NormECTight lower bounds and optimal algorithms for stochastic nonconvex optimization with heavy-tailed noise
28th International Conference on Artificial Intelligence and Statistics (AISTATS 2026)
arXiv poster
Algorithms:
NSGD-MVR, NSGD-Hess, D-clip-NSGD-MVR, Clipped NSGD-Hess, SGD-MVRMuon is provably faster with momentum variance reduction
arXiv
Algorithms:
Muon-MVR, Gluon-MVR-1, Gluon-MVR-2, Gluon-MVR-3Better LMO-based momentum methods with second-order information
arXiv
Algorithms:
LMO-SOMImproved convergence in parameter-agnostic error feedback through momentum
arXiv
Algorithms:
‖ EF21-SGDM ‖, ‖ EF21-IGT ‖, ‖ EF21-RHM ‖, ‖ EF21-HM ‖, ‖ EF21-MVR ‖Beyond the ideal: Analyzing the inexact Muon update
28th International Conference on Artificial Intelligence and Statistics (AISTATS 2026)
arXiv
Algorithms:
MuonSecond-order optimization under heavy-tailed noise: Hessian clipping and sample complexity limits
Advances in Neural Information Processing Systems 39 (NeurIPS 2025)
arXiv poster
Algorithms:
NSGDHess, Clip-NSGDHessDrop-Muon: Update less, converge faster
arXiv
Algorithms:
Drop-MuonNon-Euclidean broximal point method: a blueprint for geometry-aware optimization
arXiv
Algorithms:
BPMError feedback for Muon and friends
14th International Conference on Learning Representations (ICLR 2026)
arXiv poster
Algorithms:
EF21-MuonLocal SGD and federated averaging through the lens of time complexity
arXiv
Algorithms:
Dual Local SGD, Decaying Local SGD, Decaying Local ASGDRingleader ASGD: The first asynchronous SGD with optimal time complexity under data heterogeneity
14th International Conference on Learning Representations (ICLR 2026)
arXiv poster slides
Algorithms:
Ringleader ASGDConvergence Analysis of the ProbAbilistic Gradient Estimator Algorithm for Weakly Convex Finite-Sum Optimization
To Appear In: Journal of Optimization Theory and Applications
arXiv
Algorithms:
PAGEBernoulli-LoRA: A theoretical framework for randomized low-rank adaptation
arXiv
Algorithms:
Bernoulli-LoRAFrom Muon to Gluon: bridging theory and practice of LMO-based optimizers for LLMs
43rd International Conference on Machine Learning (ICML 2026)
arXiv slides
Algorithms:
Gluon, Muon, ScionThe stochastic multi-proximal method for nonsmooth optimization
arXiv
Algorithms:
SMPM, FedSMPM, Point-SAGA, ProxSkip, Davis-YinThanos: a block-wise pruning algorithm for efficient large language model compression
arXiv
Algorithms:
ThanosCollaborative value function estimation under model mismatch: a federated temporal difference analysis
Machine Learning and Knowledge Discovery in Databases. Research Track (ECML PKDD 2025)
arXiv
Algorithms:
FedTD (0)BurTorch: Revisiting training from first principles by coupling autodiff, math optimization, and systems
arXiv
Algorithms:
BurTorchSmoothed normalization for efficient distributed private optimization
arXiv poster
Algorithms:
α-𝖭𝗈𝗋𝗆𝖤𝖢A novel unified parametric assumption for nonconvex optimization
arXiv
Algorithms:
GD, SGDDouble momentum and error feedback for clipping with fast rates and differential privacy
arXiv
Algorithms:
Clip21-SGD2MRevisiting stochastic proximal point methods: generalized smoothness and similarity
Journal of Nonlinear and Variational Analysis 10(3):471-505, 2026
arXiv
Algorithms:
SPPMThe ball-proximal (="broximal") point method: a new algorithm, convergence theory, and applications
arXiv video slides
Algorithms:
BPM, ‖ PPM ‖ATA: Adaptive task allocation for efficient resource management in distributed machine learning
42nd International Conference on Machine Learning (ICML 2025)
arXiv poster
Algorithms:
ATASymmetric pruning of large language models
arXiv poster
Algorithms:
Symmetric Wanda, R2-DSnoTRingmaster ASGD: The first asynchronous SGD with optimal time complexity
42nd International Conference on Machine Learning (ICML 2025)
arXiv poster
Algorithms:
Naive Optimal ASGD, Ringmaster ASGD
Prepared in 2024
On the convergence of DP-SGD with adaptive clipping
arXiv
Algorithms:
QC-SGD, DP-QC-SGDMARINA-P: Superior performance in non-smooth federated optimization with adaptive stepsizes
arXiv
Algorithms:
MARINA-PDifferentially private random block coordinate descent
arXiv
Algorithms:
DP-SkGD, DP-SkGD-BS, DP-CDSpeeding up stochastic proximal optimization in the high Hessian dissimilarity setting
arXiv
Algorithms:
L-SVRPMethods with local steps and random reshuffling for generally smooth non-convex federated optimization
13th International Conference on Learning Representations (ICLR 2025)
arXiv poster
Algorithms:
Clip-LocalGDJ, CLERR, Clipped RR-CLIPushing the limits of large language model quantization via the linearity theorem
The 2025 Annual Conference of the Nations of the Americas Chapter of the ACL (NAACL 2025)
arXiv
Algorithms:
HIGGSError feedback under $(L_0,L_1)$-smoothness: normalization and momentum
Advances in Neural Information Processing Systems 39 (NeurIPS 2025)
arXiv poster
Algorithms:
‖ EF21‖, ‖ EF21-SGDM ‖Tighter performance theory of FedExProx
14th International Conference on Learning Representations (ICLR 2026)
arXiv poster
Algorithms:
FedExProxUnlocking FedNL: Self-contained compute-optimized implementation
arXiv
Algorithms:
FedNL, FedNL-LS, FedNL-PPRandomized asymmetric chain of LoRA: The first meaningful theoretical framework for low-rank adaptation
arXiv
Algorithms:
RAC-LoRA, Fed-RAC-LoRAMindFlayer SGD: Efficient parallel SGD in the presence of heterogeneous and random worker compute times
41st Conference on Uncertainty in Artificial Intelligence (UAI 2025)
NeurIPS 2024 Workshop: Optimization for Machine Learning (OPT 2024)
to be presented at Conference on the Mathematical Theory of Deep Neural Networks (DeepMath 2024) arXiv poster
Algorithms:
MindFlayer SGD, Vecna SGDOn the convergence of FedProx with extrapolation and inexact prox
NeurIPS 2024 Workshop: Optimization for Machine Learning (OPT 2024)
arXiv
Algorithms:
FedExProxMethods for convex $(L_0,L_1)$-smooth optimization: clipping, acceleration, and adaptivity
13th International Conference on Learning Representations (ICLR 2025)
arXiv poster
Algorithms:
L0L1-GD, L0L1-GD-PS, L0L1-STM, L0L1-AdGD, L0L1-SGD, L0L1-SGD-PSCohort squeeze: Beyond a single communication round per cohort in cross-device federated learning
Oral at the NeurIPS 2024 Federated Learning Workshop
arXiv
Algorithms:
SPPM-ASSparse-ProxSkip: Accelerated sparse-to-sparse training in federated learning
arXiv
Algorithms:
Sparse-ProxSkipSPAM: Stochastic proximal point method with momentum variance reduction for non-convex cross-device federated learning
NeurIPS 2024 Workshop: Optimization for Machine Learning (OPT 2024)
arXiv
Algorithms:
SPAMA simple linear convergence analysis of the Point-SAGA algorithm
arXiv
Algorithms:
Point-SAGALocal curvature descent: Squeezing more curvature out of standard and Polyak gradient descent
Advances in Neural Information Processing Systems 39 (NeurIPS 2025)
NeurIPS 2024 Workshop: Optimization for Machine Learning (OPT 2024)
arXiv poster
Algorithms:
LCD1, LCD2, LCD3On the optimal time complexities in decentralized stochastic asynchronous optimization
Advances in Neural Information Processing Systems 38 (NeurIPS 2024)
arXiv poster
Algorithms:
Fragile SGD, Amelie SGDA unified theory of stochastic proximal point methods without smoothness
arXiv
Algorithms:
SPPM, SPPM-LC, SPPM-NS, SPPM-AS, SPPM*, SPPM-GC, L-SVRP, Point SAGAMicroAdam: Accurate adaptive optimization with low space overhead and provable convergence
Advances in Neural Information Processing Systems 38 (NeurIPS 2024)
arXiv poster
Algorithms:
MicroAdamFreya PAGE: First optimal time complexity for large-scale nonconvex finite-sum optimization with heterogeneous asynchronous computations
Advances in Neural Information Processing Systems 38 (NeurIPS 2024)
arXiv poster slides
Algorithms:
Freya PAGE, Freya SGDPV-Tuning: Beyond straight-through estimation for extreme LLM compression
Advances in Neural Information Processing Systems 38 (NeurIPS 2024)
Oral at NeurIPS 2024 (0.4\% acceptance rate)
arXiv poster
Algorithms:
PVStochastic proximal point methods for monotone inclusions under expected similarity
NeurIPS 2024 Workshop: Optimization for Machine Learning (OPT 2024)
arXiv
Algorithms:
SPPM, SPPM-OC, L-SVRPThe power of extrapolation in federated learning
Advances in Neural Information Processing Systems 38 (NeurIPS 2024)
arXiv poster
Algorithms:
FedExProx, FedExProx-GraDS, FedExProx-StoPSFedComLoc: Communication-efficient distributed training of sparse and quantized models
Transactions on Machine Learning Research (TMLR 2025)
arXiv
Algorithms:
FedComLocStreamlining in the Riemannian realm: Efficient Riemannian optimization with loopless variance reduction
arXiv
Best Paper Award (runner-up), International Conference on Computational Optimization (ICOMP 2025)
Algorithms:
R-LSVRG, R-PAGE, R-MARINALoCoDL: Communication-efficient distributed learning with local training and compression
13th International Conference on Learning Representations (ICLR 2025)
NeurIPS 2024 Workshop: Optimization for Machine Learning (OPT 2024)
Spotlight at ICLR 2025
arXiv poster
Algorithms:
LoCoDLImproving the worst-case bidirectional communication complexity for nonconvex distributed optimization under function similarity
Advances in Neural Information Processing Systems 38 (NeurIPS 2024)
Spotlight at NeurIPS 2024
arXiv poster
Algorithms:
MARINA-P, M3Shadowheart SGD: Distributed asynchronous SGD with optimal time complexity under arbitrary computation and communication heterogeneity
Advances in Neural Information Processing Systems 38 (NeurIPS 2024)
arXiv poster slides
Algorithms:
Shadowheart SGDCorrelated quantization for faster nonconvex distributed optimization
41st Conference on Uncertainty in Artificial Intelligence (UAI 2025)
Oral at UAI 2025
arXiv
Algorithms:
MARINA, PermK+CQPrepared in 2023
FedP3: Personalized and privacy-friendly federated network pruning under model heterogeneity
12th International Conference on Learning Representations (ICLR 2024)
arXiv poster
Algorithms:
FedP3Error feedback reloaded: From quadratic to arithmetic mean of smoothness constants
12th International Conference on Learning Representations (ICLR 2024)
arXiv
Algorithms:
EF21-W, EF21Kimad: Adaptive gradient compression with bandwidth awareness
Proceedings of the 4th International Workshop on Distributed Machine Learning, 25--48, 2023 (DistributedML 2023)
arXiv
Algorithms:
Kimad, Kimad+Federated learning is better with non-homomorphic encryption
Proceedings of the 4th International Workshop on Distributed Machine Learning, 49--84, 2023 (DistributedML 2023)
arXiv
Algorithms:
DCGD/PermK/AESMAST: model-agnostic sparsified training
13th International Conference on Learning Representations (ICLR 2025)
arXiv
Algorithms:
double sketched (S)GD, distributed double sketched GD, L-SVRDSG, S-PAGEByzantine robustness and partial participation can be achieved simultaneously: just clip gradient differences
Advances in Neural Information Processing Systems 38 (NeurIPS 2024)
arXiv poster
Algorithms:
Byz-VR-MARINA-PPConsensus-based optimization with truncated noise
arXiv
Algorithms:
CBOCommunication compression for Byzantine robust learning: New efficient algorithms and improved rates
26th International Conference on Artificial Intelligence and Statistics (AISTATS 2024)
arXiv poster
Algorithms:
Byz-VR-MARINA, Byz-DASHA-PAGE, Byz-EF21, Byz-EF21-BCMARINA meets matrix stepsizes: Variance reduced distributed non-convex optimization
arXiv
Algorithms:
det-MARINAHigh-probability convergence for composite and distributed stochastic minimization and variational inequalities with heavy-tailed noise
41st International Conference on Machine Learning (ICML 2024)
Oral (144/9473 = top 1.5\%)
arXiv poster
Algorithms:
DProx-clipped-SGD-shift, DProx-clipped-SSTM-shiftTowards a better theoretical understanding of independent subnetwork training
41st International Conference on Machine Learning (ICML 2024)
Federated Learning and Analytics in Practice: Algorithms, Systems, Applications, and Opportunities (ICML 2023 Workshop)
arXiv poster
Algorithms:
ISTUnderstanding progressive training through the framework of randomized coordinate descent
26th International Conference on Artificial Intelligence and Statistics (AISTATS 2024)
arXiv
Algorithms:
RPTImproving accelerated federated learning with compression and importance sampling
Federated Learning and Analytics in Practice: Algorithms, Systems, Applications, and Opportunities (ICML 2023 Workshop)
arXiv
Algorithms:
5GCS-CC, 5GCS-ABClip21: Error feedback for gradient clipping
arXiv
Algorithms:
Clip21-Avg, Clip21-GD, DP-Clip21-GD, Press-Clip21-GDQuantize once, train fast: allreduce-compatible compression with provable guarantees
28th European Conference on Artificial Intelligence (ECAI 2025)
arXiv
Algorithms:
Global-QSGDA guide through the zoo of biased SGD
Advances in Neural Information Processing Systems 36 (NeurIPS 2023)
arXiv poster
Algorithms:
BiasedSGDError feedback shines when features are rare
arXiv
Algorithms:
EF21Momentum provably improves error feedback!
Advances in Neural Information Processing Systems 36 (NeurIPS 2023)
Federated Learning and Analytics in Practice: Algorithms, Systems, Applications, and Opportunities (ICML 2023 Workshop)
arXiv poster
Algorithms:
EF21-SGDM-ideal, EF21-SGDM, EF21-SGD2MExplicit personalization and local training: double communication acceleration in federated learning
Transactions on Machine Learning Research (TMLR 2025)
arXiv
Algorithms:
ScafflixOptimal time complexities of parallel stochastic optimization methods under a fixed computation model
Advances in Neural Information Processing Systems 36 (NeurIPS 2023)
arXiv video poster slides
Algorithms:
Rennala SGD, Malenia SGD2Direction: Theoretically faster distributed training with bidirectional communication compression
Advances in Neural Information Processing Systems 36 (NeurIPS 2023)
arXiv poster
Algorithms:
2DirectionDet-CGD: Compressed gradient descent with matrix stepsizes for non-convex optimization
12th International Conference on Learning Representations (ICLR 2024)
arXiv poster
Algorithms:
Det-CGDELF: Federated Langevin algorithms with primal, dual and bidirectional compression
41st Conference on Uncertainty in Artificial Intelligence (UAI 2025)
Federated Learning and Analytics in Practice: Algorithms, Systems, Applications, and Opportunities (ICML 2023 Workshop)
arXiv
Algorithms:
ELF, P-ELF, D-ELF, B-ELFTAMUNA: Doubly accelerated distributed optimization under partial participation
arXiv
Algorithms:
TAMUNAFederated learning with regularized client participation
Federated Learning and Analytics in Practice: Algorithms, Systems, Applications, and Opportunities (ICML 2023 Workshop)
arXiv
Algorithms:
RR-CLIHigh-probability bounds for stochastic optimization and variational inequalities: the case of unbounded variance
40th International Conference on Machine Learning (ICML 2023)
arXiv poster
Algorithms:
clipped-SGD, clipped-SSTM, R-clipped-SSTMCatalyst acceleration of error compensated methods leads to better communication complexity
25th International Conference on Artificial Intelligence and Statistics (AISTATS 2023)
arXiv poster
Algorithms:
ECSPDC, EC-LSVRG + Catalyst, EC-SDCA + CatalystConvergence of first-order algorithms for meta-learning with Moreau envelopes
Federated Learning and Analytics in Practice: Algorithms, Systems, Applications, and Opportunities (ICML 2023 Workshop)
arXiv
Algorithms:
FO-MuML
Prepared in 2022
Can 5th generation local training methods support client sampling? Yes!
25th International Conference on Artificial Intelligence and Statistics (AISTATS 2023)
arXiv poster
Algorithms:
5GCSAdaptive compression for communication-efficient distributed training
Transactions on Machine Learning Research (TMLR 2023)
arXiv
Algorithms:
AdaCGDA damped Newton method achieves global $O (1/k^2)$ and local quadratic convergence rate
Advances in Neural Information Processing Systems 35 (NeurIPS 2022)
arXiv poster
Algorithms:
AIC NewtonGradSkip: Communication-accelerated local gradient methods with better computational complexity
arXiv
Algorithms:
GradSkip, GradSkip+CompressedScaffnew: the first theoretical double acceleration of communication from local training and compression in distributed optimization
Optimization, 2026
arXiv
Algorithms:
CompressedScaffnewImproved Stein variational gradient descent with importance weights
arXiv
Algorithms:
beta-SVGDEF21-P and friends: Improved theoretical communication complexity for distributed optimization with bidirectional compression
40th International Conference on Machine Learning (ICML 2023)
arXiv poster
Algorithms:
EF21-P, EF21-P + DIANA, EF21-P + DCGDMinibatch stochastic three points method for unconstrained smooth minimization
38th AAAI Conference on Artificial Intelligence (AAAI 2024)
arXiv
Algorithms:
MiSTPPersonalized federated learning with communication compression
Transactions on Machine Learning Research (TMLR 2023)
arXiv github
Algorithms:
Compressed L2GDAdaptive learning rates for faster stochastic gradient methods
arXiv
Algorithms:
StoPS, GraDs, StoP, GraDRandProx: Primal-dual optimization algorithms with randomized proximal updates
11th International Conference on Learning Representations (ICLR 2023)
OPT2022: 14th Annual Workshop on Optimization for Machine Learning (NeurIPS 2022 Workshop)
arXiv video poster
Algorithms:
RandProx, RandProx-FB, RandProx-LC, RandProx-CP, RandProx-ADMM, RandProx-DYVariance reduced ProxSkip: algorithm, theory and application to federated learning
Advances in Neural Information Processing Systems 35 (NeurIPS 2022)
arXiv
Algorithms:
ProxSkip-VR, ProxSkip-GD, ProxSkip-SGD, ProxSkip-LSVRG, ProxSkip-HUBCommunication acceleration of local gradient methods via an accelerated primal-dual algorithm with inexact prox
Advances in Neural Information Processing Systems 35 (NeurIPS 2022)
arXiv
Algorithms:
APDA, APDA with Inexact Prox, APDA with Inexact Prox and Accelerated GossipShifted compression framework: generalizations and improvements
38th Conference on Uncertainty in Artificial Intelligence (UAI 2022)
arXiv poster
Algorithms:
DCGD-SHIFTA Note on the convergence of mirrored Stein variational gradient descent under (L_0, L_1) smoothness condition
arXiv
Algorithms:
MSVGDDon't compress gradients in random reshuffling: compress gradient differences
Advances in Neural Information Processing Systems 38 (NeurIPS 2024)
Federated Learning and Analytics in Practice: Algorithms, Systems, Applications, and Opportunities (ICML 2023 Workshop)
arXiv poster
Algorithms:
Q-RR, DIANA-RR, Q-NASTYA, DIANA-NASTYADistributed Newton-type methods with communication compression and Bernoulli aggregation
Transactions on Machine Learning Research (TMLR 2023)
NeurIPS Workshop 2022 (Order up! The Benefits of Higher-Order Optimization in Machine Learning)
arXiv
Algorithms:
Newton-3PC, Newton-3PC-BC, Newton-3PC-BC-PPCertified robustness in federated learning
NeurIPS Workshop 2022 (Federated Learning)
arXiv
Sharper rates and flexible framework for nonconvex SGD with client and data sampling
Transactions on Machine Learning Research (TMLR 2023)
arXiv
Algorithms:
PAGEFederated sampling with Langevin algorithm under isoperimetry
Transactions on Machine Learning Research (TMLR 2024)
arXiv
Algorithms:
Langevin-MarinaVariance reduction is an antidote to Byzantines: better rates, weaker assumptions and communication compression as a cherry on the top
11th International Conference on Learning Representations (ICLR 2023)
arXiv poster
Algorithms:
Byz-VR-MARINAConvergence of Stein variational gradient descent under a weaker smoothness condition
25th International Conference on Artificial Intelligence and Statistics (AISTATS 2023)
arXiv poster
Algorithms:
SVGDA computation and communication efficient method for distributed nonconvex problems in the partial participation setting
Advances in Neural Information Processing Systems 36 (NeurIPS 2023)
arXiv poster
Algorithms:
DASHA-PP, DASHA-PP-PAGE, DASHA-PP-FINITE-MVR, DASHA-PP-MVREF-BV: A unified theory of error feedback and variance reduction mechanisms for biased and unbiased compression in distributed optimization
Advances in Neural Information Processing Systems 35 (NeurIPS 2022)
arXiv poster
Algorithms:
EF-BVFederated random reshuffling with compression and variance reduction
arXiv
Algorithms:
FedCRR, FedCRR-VR, FedCRR-VR-2FedShuffle: Recipes for better use of local work in federated learning
Transactions on Machine Learning Research (TMLR 2022)
arXiv
Algorithms:
FedShuffleProxSkip: Yes! Local gradient steps provably lead to communication acceleration! Finally!
39th International Conference on Machine Learning (ICML 2022)
arXiv slides video
Algorithms:
ProxSkip, Scaffnew, SProxSkip, SplitSkip, Decentralized ScaffnewOptimal algorithms for decentralized stochastic variational inequalities
Advances in Neural Information Processing Systems 35 (NeurIPS 2022)
arXiv
Algorithms:
Algorithm 1, Algorithm 2DASHA: Distributed nonconvex optimization with communication compression and optimal oracle complexity
10th International Conference on Learning Representations (ICLR 2023)
Oral Paper at ICLR 2023
arXiv poster
Algorithms:
DASHA, DASHA-PAGE, DASHA-MVR3PC: Three point compressors for communication-efficient distributed training and a better theory for lazy aggregation
39th International Conference on Machine Learning (ICML 2022)
arXiv poster
Algorithms:
3PC, LAG, CLAG, EF21BEER: Fast $O (1/T)$ rate for decentralized nonconvex optimization with communication compression
Advances in Neural Information Processing Systems 35 (NeurIPS 2022)
arXiv poster
Algorithms:
BEERServer-side stepsizes and sampling without replacement provably help in federated optimization
Proceedings of the 4th International Workshop on Distributed Machine Learning, 85--104, 2023 (DistributedML 2023) arXiv
Algorithms:
Nastya
Prepared in 2021
Accelerated primal-dual gradient method for smooth and convex-concave saddle-point problems with bilinear coupling
Advances in Neural Information Processing Systems 35 (NeurIPS 2022)
arXiv
Algorithms:
APDGFaster rates for compressed federated learning with client-variance reduction
SIAM Journal on Mathematics of Data Science 6 (1):154-175, 2024
arXiv
Algorithms:
COFIG, FRECONFL_PyTorch: optimization research simulator for federated learning
Proceedings of the 2nd ACM International Workshop on Distributed Machine Learning
arXiv
Algorithms:
FL_PyTorchFLIX: A simple and communication-efficient alternative to local methods in federated learning
24th International Conference on Artificial Intelligence and Statistics (AISTATS 2022)
arXiv poster
Algorithms:
FLIXBasis matters: better communication-efficient second order methods for federated learning
24th International Conference on Artificial Intelligence and Statistics (AISTATS 2022)
arXiv poster
Algorithms:
BL1, BL2, BL3Distributed methods with compressed communication for solving variational inequalities, with theoretical guarantees
Advances in Neural Information Processing Systems 35 (NeurIPS 2022)
arXiv
Algorithms:
MASHA1, MASHA2Permutation compressors for provably faster distributed nonconvex optimization
10th International Conference on Learning Representations (ICLR 2022)
arXiv video 1 video 2 video 3 poster
Algorithms:
MARINAEF21 with bells & whistles: practical algorithmic extensions of modern error feedback
Journal of Machine Learning Research, 2025
arXiv github
Algorithms:
EF21-SGD, EF21-PAGE, EF21-PP, EF21-BC, EF21-HB, EF21-ProxError compensated loopless SVRG, Quartz, and SDCA for distributed optimization
arXiv
Algorithms:
EC-LSVRG, EC-SDCA, EC-QuartzDoubly adaptive scaled algorithm for machine learning using second-order information
10th International Conference on Learning Representations (ICLR 2022)
arXiv poster
Algorithms:
OASISFedPAGE: A fast local stochastic gradient method for communication-efficient federated learning
arXiv
Algorithms:
FedPAGECANITA: Faster rates for distributed convex optimization with communication compression
Advances in Neural Information Processing Systems 34 (NeurIPS 2021)
arXiv poster
Algorithms:
CANITAEF21: A new, simpler, theoretically better, and practically faster error feedback
Advances in Neural Information Processing Systems 34 (NeurIPS 2021)
NeurIPS 2021 oral paper (less than 1\% acceptance rate)
arXiv slides video 1 video 2 poster github
Algorithms:
EF21, EF21+Lower bounds and optimal algorithms for smooth and strongly convex decentralized optimization over time-varying networks
Advances in Neural Information Processing Systems 34 (NeurIPS 2021)
arXiv poster
Algorithms:
ADOM+Theoretically better and numerically faster distributed optimization with smoothness-aware quantization techniques
Advances in Neural Information Processing Systems 35 (NeurIPS 2022)
arXiv poster
Algorithms:
DCGD+, DIANA+A convergence theory for SVGD in the population limit under Talagrand’s inequality T1
39th International Conference on Machine Learning (ICML 2022)
arXiv
Algorithms:
SVGDMURANA: A generic framework for stochastic variance-reduced optimization
Mathematical and Scientific Machine Learning 2022 (MSML 2022)
arXiv
Algorithms:
MURANA, ELVIRAFedNL: Making Newton-type methods applicable to federated learning
39th International Conference on Machine Learning (ICML 2022)
arXiv poster
Algorithms:
FedNL, FedNL-PP, FedNL-CR, FedNL-LS, FedNL-BC, N0, NSRandom reshuffling with variance reduction: new analysis and better rates
39th Conference on Uncertainty in Artificial Intelligence (UAI 2023)
arXiv video
Algorithms:
RR-SVRG, SO-SVRG, Cyclic-SVRGZeroSARAH: Efficient nonconvex finite-sum optimization with zero full gradient computation
arXiv
Algorithms:
Zero-SARAHAn optimal algorithm for strongly convex minimization under affine constraints
24th International Conference on Artificial Intelligence and Statistics (AISTATS 2022)
arXiv poster
Algorithms:
accelerated PAPCAI-SARAH: Adaptive and implicit stochastic recursive gradient methods
Transactions on Machine Learning Research (TMLR 2023)
arXiv
Algorithms:
AI-SARAHADOM: Accelerated decentralized optimization method for time-varying networks
38th International Conference on Machine Learning (ICML 2021)
NSF-TRIPODS Workshop: Communication Efficient Distributed Optimization
arXiv video poster
Algorithms:
ADOMIntSGD: Floatless compression of stochastic gradients
10th International Conference on Learning Representations (ICLR 2022)
ICLR 2022 Spotlight paper
arXiv video poster
Algorithms:
IntSGD, IntDIANAMARINA: faster non-convex distributed learning with compression
38th International Conference on Machine Learning (ICML 2021)
NSF-TRIPODS Workshop: Communication Efficient Distributed Optimization
arXiv video 1 video 2 poster
Algorithms:
MARINA, VR-MARINA, PP-MARINASmoothness matrices beat smoothness constants: better communication compression techniques for distributed optimization
Advances in Neural Information Processing Systems 34 (NeurIPS 2021)
ICLR Workshop: Distributed and Private Machine Learning
NSF-TRIPODS Workshop: Communication Efficient Distributed Optimization
arXiv video poster
Algorithms:
DCGD+, DIANA+, ADIANA+Distributed second order methods with fast rates and compressed communication
38th International Conference on Machine Learning (ICML 2021)
NSF-TRIPODS Workshop: Communication Efficient Distributed Optimization
arXiv slides video 1 video 2 video 3 poster
Algorithms:
NS, MN, NL1, NL2, CNLProximal and federated random reshuffling
39th International Conference on Machine Learning (ICML 2022)
NSF-TRIPODS Workshop: Communication Efficient Distributed Optimization
arXiv video
Algorithms:
ProxRR, FedRR
Prepared in 2020
Hyperparameter transfer learning with adaptive complexity
The 24th International Conference on Artificial Intelligence and Statistics (AISTATS 2021)
arXiv poster
Algorithms:
ABRACError compensated loopless SVRG for distributed optimization
OPT2020: 12th Annual Workshop on Optimization for Machine Learning (NeurIPS 2020 Workshop)
arXiv poster
Algorithms:
EC-LSVRGError compensated proximal SGD and RDA
OPT2020: 12th Annual Workshop on Optimization for Machine Learning (NeurIPS 2020 Workshop)
poster
Algorithms:
EC-SGD, EC-RDALocal SGD: unified theory and new efficient methods
The 24th International Conference on Artificial Intelligence and Statistics (AISTATS 2021)
arXiv video poster
Algorithms:
S-Local-SVRGA linearly convergent algorithm for decentralized optimization: sending less bits for free!
The 24th International Conference on Artificial Intelligence and Statistics (AISTATS 2021)
arXiv video poster
Optimal client sampling for federated learning
Transactions on Machine Learning Research (TMLR 2022)
Privacy Preserving Machine Learning (NeurIPS 2020 Workshop)
arXiv
Algorithms:
OCS, AOCSLinearly converging error compensated SGD
Advances in Neural Information Processing Systems 33 (NeurIPS 2020)
arXiv video poster
Algorithms:
EC-SGD-DIANA, EC-LSVRG-DIANA, EC-LSVRGstar, ...Optimal gradient compression for distributed and federated learning
SpicyFL 2020: NeurIPS Workshop on Scalability, Privacy, and Security in Federated Learning
arXiv video poster
Lower bounds and optimal algorithms for personalized federated learning
Advances in Neural Information Processing Systems 33 (NeurIPS 2020)
arXiv video
Algorithms:
APGD1, APGD2, IAPGD, AL2SGD+Distributed proximal splitting algorithms with rates and acceleration
Frontiers in Signal Processing, section Signal Processing for Communications, 2022
OPT2020: 12th Annual Workshop on Optimization for Machine Learning (NeurIPS 2020 Workshop)
Spotlight Talk
arXiv poster
Algorithms:
PD3O, PDDY, distributed PD3O, distributed PDDYVariance-reduced methods for machine learning
Proceedings of the IEEE 108 (11):1968--1983, 2020
arXiv
Algorithms:
SAG, SAGA, SVRG, SDCAError compensated distributed SGD can be accelerated
Advances in Neural Information Processing Systems 34 (NeurIPS 2021)
OPT2020: 12th Annual Workshop on Optimization for Machine Learning (NeurIPS 2020 Workshop)
arXiv poster
Algorithms:
ECLKQuasi-Newton methods for deep learning: forget the past, just sample
Optimization Methods and Software 37(5):1668-1704, 2022
2022 Charles Broyden Prize
arXiv
Algorithms:
S-LBFGS, S-LSR1PAGE: A simple and optimal probabilistic gradient estimator for nonconvex optimization
38th International Conference on Machine Learning (ICML 2021)
OPT2020: 12th Annual Workshop on Optimization for Machine Learning (NeurIPS 2020 Workshop) (Spotlight Talk)
arXiv video poster
Algorithms:
PAGEOptimal and practical algorithms for smooth and strongly convex decentralized optimization
Advances in Neural Information Processing Systems 33 (NeurIPS 2020)
arXiv
Algorithms:
APAPC, OPAPC, Algorithm 3Unified analysis of stochastic gradient methods for composite convex and smooth optimization
Journal of Optimization Theory and Applications 199:499-540, 2023
arXiv
Algorithms:
SGDA better alternative to error feedback for communication-efficient distributed learning
9th International Conference on Learning Representations (ICLR 2021)
SpicyFL 2020: NeurIPS Workshop on Scalability, Privacy, and Security in Federated Learning
The Best Paper Award at NeurIPS-20 Workshop on Scalability, Privacy, and Security in Federated Learning
arXiv poster
Algorithms:
DCSGDPrimal dual interpretation of the proximal stochastic gradient Langevin algorithm
Advances in Neural Information Processing Systems 33 (NeurIPS 2020)
arXiv
Algorithms:
PGSLAA unified analysis of stochastic gradient methods for nonconvex federated optimization
SpicyFL 2020: NeurIPS Workshop on Scalability, Privacy, and Security in Federated Learning
arXiv video
Algorithms:
DC-GD, DC-SGD, DC-LSVRG, DC-SAGA, DIANA-GD, DIANA-SGD, DIANA-LSVRG, DIANA-SAGARandom reshuffling: simple analysis with vast improvements
Advances in Neural Information Processing Systems 33 (NeurIPS 2020)
arXiv video poster code
Algorithms:
RR, SO, IGAdaptive learning of the optimal mini-batch size of SGD
OPT2020: 12th Annual Workshop on Optimization for Machine Learning (NeurIPS 2020 Workshop)
arXiv poster
Algorithms:
SGD with Adaptive Batch sizeDualize, split, randomize: fast nonsmooth optimization algorithms
Journal of Optimization Theory and Applications 195: 102-130, 2022
OPT2020: 12th Annual Workshop on Optimization for Machine Learning (NeurIPS 2020 Workshop)
arXiv poster
Algorithms:
PDDY, SPDDY, SPD3O, SPAPCOn the convergence analysis of asynchronous SGD for solving consistent linear systems
Linear Algebra and its Applications, 2022
arXiv
Algorithms:
DASGDFrom local SGD to local fixed point methods for federated learning
37th International Conference on Machine Learning (ICML 2020)
arXiv video
Algorithms:
LDFPM, RDFPMOn biased compression for distributed learning
Accepted to Journal of Machine Learning Research, 2022
SpicyFL 2020: NeurIPS Workshop on Scalability, Privacy, and Security in Federated Learning
arXiv poster
Algorithms:
CGD, Distributed SGD with Error FeedbackAcceleration for compressed gradient descent in distributed and federated optimization
37th International Conference on Machine Learning (ICML 2020)
arXiv
Algorithms:
ACGD, ADIANAFast linear convergence of randomized BFGS
arXiv
Algorithms:
RBFGSStochastic subspace cubic Newton method
37th International Conference on Machine Learning (ICML 2020)
arXiv
Algorithms:
SSCNUncertainty principle for communication compression in distributed and federated learning and the search for an optimal compressor
Information and Inference: A Journal of the IMA, 1--24, 2021
arXiv
Algorithms:
KCFederated learning of a mixture of global and local models
SpicyFL 2020: NeurIPS Workshop on Scalability, Privacy, and Security in Federated Learning
arXiv slides video poster
Algorithms:
L2GD, L2SGD+Adaptivity of stochastic gradient methods for nonconvex optimization
OPT2020: 12th Annual Workshop on Optimization for Machine Learning (NeurIPS 2020 Workshop)
SIAM Journal on Mathematics of Data Science 4(2):634--648, 2022
arXiv poster
Algorithms:
Geometrized SARAHVariance reduced coordinate descent with acceleration: new method with a surprising application to finite-sum problems
37th International Conference on Machine Learning (ICML 2020)
arXiv
Algorithms:
ASVRCDBetter theory for SGD in the nonconvex world
Transactions on Machine Learning Research (TMLR 2022)
arXiv
Algorithms:
SGDPrepared in 2019
Tighter theory for local SGD on identical and heterogeneous data
The 23rd International Conference on Artificial Intelligence and Statistics (AISTATS 2020)
arXiv
Algorithms:
Local SGDDistributed fixed point methods with compressed iterates
arXiv preprint
Algorithms:
FPMCI, VR-FPMCIIntML: Natural compression for distributed deep learning
Workshop on AI Systems at Symposium on Operating Systems Principles 2019 (SOSP'19)
arXiv pdf
Algorithms:
NC, natural ditheringStochastic Newton and cubic Newton methods with simple local linear-quadratic rates
NeurIPS 2019 Workshop Beyond First Order Methods in ML
arXiv poster
Algorithms:
SN, SCNBetter communication complexity for local SGD
NeurIPS 2019 Workshop on Federated Learning for Data Privacy and Confidentiality
arXiv poster
Algorithms:
local SGDGradient descent with compressed iterates
NeurIPS 2019 Workshop on Federated Learning for Data Privacy and Confidentiality
arXiv poster
Algorithms:
GDCIFirst analysis of local GD on heterogeneous data
NeurIPS 2019 Workshop on Federated Learning for Data Privacy and Confidentiality
arXiv
Algorithms:
local GDStochastic convolutional sparse coding
International Symposium on Vision, Modeling and Visualization 2019
VMV Best Paper Award, 2019 link
arXiv
Algorithms:
SBCSC, SOCSCL-SVRG and L-Katyusha with arbitrary sampling
Journal of Machine Learning Research 22(112):1−47, 2021
arXiv video
Algorithms:
L-SVRG, L-KatyushaMISO is making a comeback with better proofs and rates
arXiv
Algorithms:
MISOA stochastic derivative free optimization method with momentum
8th International Conference on Learning Representations (ICLR 2020)
arXiv poster
Algorithms:
SMTPStochastic Sign Descent Methods: New Algorithms and Better Theory
38th International Conference on Machine Learning (ICML 2021)
OPT2020: 12th Annual Workshop on Optimization for Machine Learning (NeurIPS 2020 Workshop)
arXiv poster
Algorithms:
signSGD, signSGDmajStochastic proximal Langevin algorithm: potential splitting and nonasymptotic rates
33rd Conference on Neural Information Processing Systems (NeurIPS 2019)
arXiv poster
Algorithms:
SPLADirect nonlinear acceleration
EURO Journal on Computational Optimization 10, 2022, 100047
arXiv
Algorithms:
DNAA stochastic decoupling method for minimizing the sum of smooth and non-smooth functions
arXiv
Algorithms:
SDMRevisiting stochastic extragradient
The 23rd International Conference on Artificial Intelligence and Statistics (AISTATS 2020)
NeuriPS 2019 Workshop on Smooth Games Optimization and Machine Learning
arXiv
Algorithms:
stochastic extragradientOne method to rule them all: variance reduction for data, parameters and many new methods
arXiv poster
Algorithms:
GJS + 17 algorithmsA unified theory of SGD: variance reduction, sampling, quantization and coordinate descent
The 23rd International Conference on Artificial Intelligence and Statistics (AISTATS 2020)
arXiv
Algorithms:
SGD-MB, SGD-star, N-SAGA, N-SEGA, Q-SGD-SRNatural compression for distributed deep learning
Mathematical and Scientific Machine Learning 2022 (MSML 2022)
arXiv poster
Algorithms:
NC, natural ditheringRSN: Randomized Subspace Newton
33rd Conference on Neural Information Processing Systems (NeurIPS 2019)
arXiv poster
Algorithms:
RSNBest pair formulation & accelerated scheme for non-convex principal component pursuit
IEEE Transactions on Signal Processing 68:6128-6141, 2020
arXiv
Algorithms:
accelerated proximal gradientRevisiting randomized gossip algorithms: general framework, convergence rates and novel block and accelerated protocols
IEEE Transactions on Information Theory 67(12):8300--8324, 2021
arXiv
Algorithms:
block gossip, accelerated gossip, dual gossipConvergence analysis of inexact randomized iterative methods
SIAM Journal on Scientific Computing 42(6), A3979–A4016, 2020
arXiv
Algorithms:
iBasic, iSDSA, iSGD, iSPM, iRBK, iRBCDScaling distributed machine learning with in-network aggregation
The 18th USENIX Symposium on Networked Systems Design and Implementation (NSDI '21 Fall)
arXiv
Algorithms:
SwitchMLStochastic distributed learning with gradient quantization and double variance reduction
Optimization Methods and Software 38(1):91-106, 2023
2023 Charles Broyden Prize
arXiv
Algorithms:
DIANA, VR-DIANA, SVRG-DIANAStochastic three points method for unconstrained smooth minimization
SIAM Journal on Optimization 30(4):2726-2749, 2020
arXiv
Algorithms:
STPA stochastic derivative-free optimization method with importance sampling
34th AAAI Conference on Artificial Intelligence (AAAI 2020)
arXiv poster
Algorithms:
STP_IS99% of distributed optimization is a waste of time: the issue and how to fix it
36th Conference on Uncertainty in Artificial Intelligence (UAI 2020)
arXiv
Algorithms:
IBCD, ISAGA, ISGD, IASGD, ISEGADistributed learning with compressed gradient differences
Optimization Methods and Software 40(5):1181--1196, 2025
arXiv
Algorithms:
DIANASGD: general analysis and improved rates
Proceedings of the 36th International Conference on Machine Learning, PMLR 97:5200-5209, 2019
arXiv video poster
Algorithms:
SGD-ASDon’t jump through hoops and remove those loops: SVRG and Katyusha are better without the outer loop
31st International Conference on Learning Theory (ALT 2020)
arXiv
Algorithms:
L-SVRG, L-KatyushaSAGA with arbitrary sampling
Proceedings of the 36th International Conference on Machine Learning, PMLR 97:5190-5199, 2019
arXiv poster
Algorithms:
SAGA-AS
Prepared in 2018
New convergence aspects of stochastic gradient algorithms
Journal of Machine Learning Research 20(176):1-49, 2019
arXiv
Algorithms:
SGD, Hogwild!A privacy preserving randomized gossip algorithm via controlled noise insertion
NeurIPS Privacy Preserving Machine Learning Workshop, 2018
arXiv poster
Algorithms:
Private Gossip with Controlled Noise InsertionA stochastic penalty model for convex and nonconvex optimization with big constraints
arXiv poster
Algorithms:
SGD, Increasing Penalty MethodProvably accelerated randomized gossip algorithms
2019 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP 2019)
arXiv
Algorithms:
AccGossipAccelerated coordinate descent with arbitrary sampling and best rates for minibatches
22nd International Conference on Artificial Intelligence and Statistics (2019
arXiv poster
Algorithms:
ACDNonconvex variance reduced optimization with arbitrary sampling
Proceedings of the 36th International Conference on Machine Learning, PMLR 97:2781-2789, 2019
Horváth: Best DS3 Poster Award, Paris, 2018 (link)
arXiv poster
Algorithms:
SVRG, SAGA, SARAHSEGA: Variance reduction via gradient sketching
Advances in Neural Information Processing Systems 31:2082-2093, 2018
arXiv slides video poster
Algorithms:
SEGAAccelerated Bregman proximal gradient methods for relatively smooth convex optimization
Computational Optimization and Applications 79:405–440, 2021
arXiv
Algorithms:
ABPG, ABDAMatrix completion under interval uncertainty: highlights
Lecture Notes in Computer Science, ECML-PKDD 2018
Algorithms:
MACOAccelerated gossip via stochastic heavy ball method
56th Annual Allerton Conference on Communication, Control, and Computing, 927-934, 2018
Press coverage [KAUST Discovery]
arXiv poster
Algorithms:
SHB, mRK, mRBKImproving SAGA via a probabilistic interpolation with gradient descent
arXiv
Algorithms:
SAGDA nonconvex projection method for robust PCA
33rd AAAI Conference on Artificial Intelligence (AAAI 2019)
arXiv
Algorithms:
alternating projection for RPCA, alternating projection for RMCStochastic quasi-gradient methods: variance reduction via Jacobian sketching
Mathematical Programming 188:135–192, 2021
arXiv slides video
Algorithms:
JacSketchWeighted low-rank approximation of matrices and background modeling
arXiv
Algorithms:
WLR, inWLRFastest rates for stochastic mirror descent methods
Computational Optimization and Applications 79:717–766, 2021
arXiv
Algorithms:
relRCD, relSGDSGD and Hogwild! convergence without the bounded gradients assumption
Proceedings of The 35th International Conference on Machine Learning, PMLR 80:3750-3758, 2018
arXiv poster
Algorithms:
SGD, Hogwild!Accelerated stochastic matrix inversion: general theory and speeding up BFGS rules for faster second-order optimization
Advances in Neural Information Processing Systems 31:1619-1629, 2018
arXiv poster
Algorithms:
ABFGSRandomized block cubic Newton method
Proceedings of The 35th International Conference on Machine Learning, PMLR 80:1290-1298, 2018
Doikov: Best Talk Award, "Control, Information and Optimization", Voronovo, Russia, 2018
arXiv poster bib
Algorithms:
RBCNStochastic spectral and conjugate descent methods
32nd Conference on Neural Information Processing Systems (NeurIPS 2018)
arXiv poster
Algorithms:
SSD, SconD, SSCD, mSSCD, iSconD, iSSDA randomized exchange algorithm for computing optimal approximate designs of experiments
Journal of the American Statistical Association, 2020
arXiv
Algorithms:
REX, OD_REX, MVEE_REXRandomized projection methods for convex feasibility problems: conditioning and convergence rates
SIAM Journal on Optimization 29(4):2814–2852, 2019
arXiv slides
Algorithms:
SPA, SAP, AvPPrepared in 2017
Momentum and stochastic momentum for stochastic gradient, Newton, proximal point and subspace descent methods
Computational Optimization and Applications 77(3):653-710, 2020
arXiv
Algorithms:
mSGD, mSN, mSPP, mSDSA, smSGD, smSN, smSPPOnline and batch supervised background estimation via L1 regression
IEEE Winter Conference on Applications in Computer Vision, 2019
arXiv
Algorithms:
IRLS, Homotopy, SGD 1, SGD 2, ALMLinearly convergent stochastic heavy ball method for minimizing generalization error
NIPS Workshop on Optimization for Machine Learning, 2017
arXiv poster
Algorithms:
SHBGlobal convergence of arbitrary-block gradient methods for generalized Polyak-Łojasiewicz functions
arXiv
The complexity of primal-dual fixed point methods for ridge regression
Linear Algebra and its Applications 556:342-372, 2018
arXiv
Algorithms:
PDFP1, PDFP2, Quartz, New Quartz, Modified QuartzFaster PET reconstruction with a stochastic primal-dual hybrid gradient method
Proceedings of SPIE, Wavelets and Sparsity XVII, Volume 10394, pages 1039410-1 - 1039410-11, 2017
pdf video poster
Algorithms:
SPDHGA batch-incremental video background estimation model using weighted low-rank approximation of matrices
IEEE International Conference on Computer Vision (ICCV) Workshops, 2017
arXiv
Algorithms:
inWLRPrivacy preserving randomized gossip algorithms
arXiv slides
Algorithms:
Private Gossip with Binary Oracle, Private Gossip with ε-Gap Oracle, Private Gossip with Controlled Noise InsertionStochastic primal-dual hybrid gradient algorithm with arbitrary sampling and imaging applications
SIAM Journal on Optimization 28(4):2783-2808, 2018
arXiv slides video poster
Algorithms:
SPDHGStochastic reformulations of linear systems: algorithms and convergence theory
SIAM Journal on Matrix Analysis and Applications 41(2):487–524, 2020
arXiv slides
Algorithms:
basic, parallel and accelerated methodsParallel stochastic Newton method
Journal of Computational Mathematics 36(3):404-425, 2018
arXiv
Algorithms:
PSNM
Prepared in 2016
Linearly convergent randomized iterative methods for computing the pseudoinverse
arXiv
Algorithms:
SATAX, SAXASRandomized distributed mean estimation: accuracy vs communication
Frontiers in Applied Mathematics and Statistics 2018
arXiv
Algorithms:
variable-size encoder, fixed-size encoderFederated learning: strategies for improving communication efficiency
NIPS Private Multi-Party Machine Learning Workshop, 2016
link [selected press coverage: The Verge - Quartz - Vice CBR - Android Authority]
arXiv poster
Algorithms:
structured updates, sketched updatesFederated optimization: distributed machine learning for on-device intelligence
link [selected press coverage: The Verge - Quartz - Vice CBR - Android Authority]
arXiv
Algorithms:
FSVRGA new perspective on randomized gossip algorithms
IEEE Global Conference on Signal and Information Processing (GlobalSIP), 440-444, 2016
arXiv poster
Algorithms:
SDA, RBK, RNMAIDE: fast and communication efficient distributed optimization
arXiv poster
Algorithms:
Inexact DANE, AIDECoordinate descent face-off: primal or dual?
Proceedings of Algorithmic Learning Theory, PMLR 83:246-267, 2018
arXiv bib
Algorithms:
NSync, QUARTZOptimization in high dimensions via accelerated, parallel and proximal coordinate descent
SIAM Review 58(4):739-771, 2016
SIAM SIGEST Award
arXiv
Algorithms:
APPROXStochastic block BFGS: squeezing more curvature out of data
Proceedings of the 33rd International Conference on Machine Learning, PMLR 48:1869-1878, 2016
arXiv poster bib
Algorithms:
Stochastic Block BFGSImportance sampling for minibatches
Journal of Machine Learning Research 19(27):1-21, 2018
arXiv bib
Algorithms:
dfSDCARandomized quasi-Newton updates are linearly convergent matrix inversion algorithms
SIAM Journal on Matrix Analysis and Applications 38(4):1380-1409, 2017
Most Downloaded SIMAX Paper (6th place: 2018)
arXiv
Algorithms:
SIMI, RBFGS, AdaRBFGS, ...
Prepared in 2015
[43] Zeyuan Allen-Zhu, Zheng Qu, Peter Richtárik and Yang Yuan
Even faster accelerated coordinate descent using non-uniform sampling
Proceedings of the 33rd International Conference on Machine Learning, PMLR 48:1110-1119, 2016
arXiv bib
Algorithms: NU_ACDM
[42] Robert M. Gower and Peter Richtárik
Stochastic dual ascent for solving linear systems
arXiv video
Algorithms: SDA
[41] Chenxin Ma, Jakub Konečný, Martin Jaggi, Virginia
Smith, Michael I Jordan, P. Richtárik and Martin Takáč
Distributed optimization with arbitrary local solvers
Optimization Methods and Software 32(4):813-848, 2017
Most-Read Paper, Optimization Methods and Software, 2017
arXiv
Algorithms: CoCoA+
[40] Martin Takáč, Peter Richtárik and Nathan Srebro
Distributed mini-batch SDCA
To appear in: Journal of Machine Learning Research
arXiv
Algorithms: mSDCA
[39] Robert M. Gower and Peter Richtárik
Randomized iterative methods for linear systems
SIAM
Journal on Matrix Analysis and Applications 36(4):1660-1690, 2015
Most Downloaded SIMAX Paper (1st place: 2017-2020)
Gower: 18th IMA Leslie Fox Prize (2nd Prize), 2017
link
arXiv
slides
Algorithms: sketch-and-project, GK, Gauss-LS, Gauss-pd
[38] Dominik Csiba and Peter Richtárik
Primal method for ERM with flexible mini-batching schemes and non-convex losses
arXiv
Algorithms: dfSDCA
[37] Jakub Konečný, Jie Liu, Peter Richtárik and Martin
Takáč
Mini-batch semi-stochastic gradient descent in the
proximal setting
IEEE
Journal of Selected Topics in Signal Processing 10(2): 242-255,
2016
arXiv
Algorithms: mS2GD
[36] Rachael Tappenden, Martin Takáč and Peter Richtárik
On the complexity of parallel coordinate descent
Optimization Methods and Software 33(2):372-395, 2018
arXiv
Algorithms: PCDM, PCDM-M
[35] Dominik Csiba, Zheng Qu and Peter Richtárik
Stochastic dual coordinate ascent with adaptive
probabilities
Proceedings
of the 32nd International Conference on Machine Learning,
PMLR 37:674-683, 2015
Csiba: Best Contribution Award (2nd
Place), Optimization and Big Data 2015
Implemented in Tensor Flow
arXiv poster bib
Algorithms: AdaSDCA and AdaSDCA+
[34] Chenxin Ma, Virginia Smith, Martin Jaggi, Michael I.
Jordan, Peter Richtárik and Martin Takáč
Adding vs. averaging in distributed primal-dual
optimization
Proceedings of the 32nd International Conference on Machine Learning,
PMLR 37:1973-1982, 2015
Smith: 2015 MLconf Industry Impact
Student Research Award link
CoCoA+ is now the default linear optimizer in Tensor Flow link
arXiv poster bib
Algorithms: CoCoA+
[33] Zheng Qu, Peter Richtárik, Martin Takáč and Olivier
Fercoq
SDNA: Stochastic dual Newton ascent for empirical risk
minimization
Proceedings
of the 33rd International Conference on Machine Learning, PMLR 48:1823-1832, 2016
arXiv slides poster bib
Algorithms: SDNA
Prepared in 2014
[32] Zheng Qu and Peter Richtárik
Coordinate descent with arbitrary sampling II: expected
separable overapproximation
Optimization
Methods and Software 31(5):858-884, 2016
arXiv
[31] Zheng Qu and Peter Richtárik
Coordinate descent with arbitrary sampling I: algorithms
and complexity
Optimization
Methods and Software 31(5):829-857, 2016
arXiv
Algorithms: ALPHA
[30] Jakub Konečný, Zheng Qu and Peter Richtárik
Semi-stochastic coordinate descent
Optimization
Methods and Software 32(5):993-1005, 2017
arXiv
Algorithms: S2CD
[29] Zheng Qu, Peter Richtárik and Tong Zhang
Quartz: Randomized dual coordinate ascent with arbitrary
sampling
Advances
in Neural Information Processing Systems 28:865-873, 2015
arXiv slides video
Algorithms: QUARTZ
[28] Jakub Konečný, Jie Liu, Peter Richtárik and Martin
Takáč
mS2GD: Mini-batch semi-stochastic gradient descent in the
proximal setting
NIPS
Workshop on Optimization for Machine Learning, 2014
arXiv poster
Algorithms: mS2GD
[27] Jakub Konečný, Zheng Qu and Peter Richtárik
S2CD: Semi-stochastic coordinate descent
NIPS
Workshop on Optimization for Machine Learning, 2014
pdf poster
Algorithms: S2CD
[26] Jakub Konečný and Peter Richtárik
Simple complexity analysis of simplified direct search
arXiv slides in Slovak
Algorithms: SDS
[25] Jakub Mareček, Peter Richtárik and Martin Takáč
Distributed block coordinate descent for minimizing
partially separable functions
Numerical
Analysis and Optimization, Springer Proceedings in Math. and
Statistics 134:261-288, 2015
arXiv
Algorithms: Distributed BCD
[24] Olivier Fercoq, Zheng Qu, Peter Richtárik and Martin
Takáč
Fast distributed coordinate descent for minimizing
non-strongly convex losses
2014
IEEE International Workshop on Machine Learning for Signal
Processing (MLSP), 2014
arXiv poster
Algorithms: Hydra^2
[23] Duncan Forgan and Peter Richtárik
On optimal solutions to planetesimal growth models
Technical Report ERGO 14-002, 2014
pdf
[22] Jakub Mareček, Peter Richtárik and Martin Takáč
Matrix completion under interval uncertainty
European
Journal of Operational Research 256(1):35-42, 2017
arXiv
Algorithms: MACO
Prepared in 2013
[21] Olivier Fercoq and Peter Richtárik
Accelerated, Parallel and PROXimal coordinate descent
SIAM
Journal on Optimization 25(4):1997-2023, 2015
Fercoq: 17th IMA Leslie Fox Prize
(Second Prize), 2015
2nd Most Downloaded SIOPT Paper (Aug
2016 - now)
arXiv video poster
Algorithms: APPROX
[20] Jakub Konečný and Peter Richtárik
Semi-stochastic gradient descent methods
Frontiers
in Applied Mathematics and Statistics 3:9, 2017
arXiv slides poster
Algorithms: S2GD and S2GD+
[19] Peter Richtárik and Martin Takáč
On optimal probabilities in stochastic coordinate descent
methods
Optimization Letters
10(6):1233-1243, 2016
arXiv poster
Algorithms: NSync
[18] Peter Richtárik and Martin Takáč
Distributed coordinate descent method for learning with
big data
Journal
of Machine Learning Research 17(75):1-25, 2016
arXiv poster
Algorithms: Hydra
[17] Olivier Fercoq and Peter Richtárik
Smooth minimization of nonsmooth functions with parallel
coordinate descent methods
Springer Proceedings in Mathematics and Statistics 279:57-96, 2019
arXiv
Algorithms: SPCDM
[16] Rachael Tappenden, Peter Richtárik and Burak Buke
Separable approximations and decomposition methods for
the augmented Lagrangian
Optimization
Methods and Software 30(3):643-668, 2015
arXiv
Algorithms: DQAM, PCDM
[15] Rachael Tappenden, Peter Richtárik and Jacek Gondzio
Inexact coordinate descent: complexity and
preconditioning
Journal
of Optimization Theory and Applications 170(1):144-176, 2016
arXiv poster
Algorithms: ICD
[14] Martin Takáč, Selin Damla Ahipasaoglu, Ngai-Man Cheung
and Peter Richtárik
TOP-SPIN: TOPic discovery via Sparse Principal component
INterference
Springer Proceedings in Mathematics and Statistics 279:157-180,
2019
arXiv poster
Algorithms: TOP-SPIN
[13] Martin Takáč, Avleen Bijral, Peter Richtárik and
Nathan Srebro
Mini-batch primal and dual methods for SVMs
Proceedings
of the 30th International Conference on Machine Learning,
2013
arXiv poster
Algorithms: minibatch SDCA and minibatch Pegasos
Prepared in 2012 or earlier
[12] Peter Richtárik, Majid Jahani, Martin Takáč and Selin Damla Ahipasaoglu
Alternating maximization: unifying framework for 8 sparse PCA formulations and efficient parallel codes
Optimization and Engineering 22:1493--1519, 2021
arXiv
Algorithms: 24am
[11] William Hulme, Peter Richtárik, Lynne McGuire and Alison Green
Optimal diagnostic tests for sporadic Creutzfeldt-Jakob disease based on SVM classification of RT-QuIC data
Technical Report, 2012
arXiv
[10] Peter Richtárik and Martin Takáč
Parallel coordinate descent methods for big data optimization
Mathematical Programming 156(1):433-484, 2016
Takáč: 16th IMA Leslie Fox Prize
(2nd Prize), 2013 link
#1 Top Trending Article in
Mathematical Programming Ser A and B (2017) link
arXiv slides
video
Algorithms: PCDM, AC/DC
[9] Peter Richtárik and Martin Takáč
Efficient serial and parallel coordinate descent methods for huge-scale truss topology design
Operations Research Proceedings 2011:27-32, Springer-Verlag, 2012
Optimization Online poster
Algorithms: Serial CD, Parallel CD
[8] Peter Richtárik and Martin Takáč
Iteration complexity of randomized block-coordinate descent methods for minimizing a composite function
Mathematical Programming 144(2):1-38, 2014
Best Student Paper (runner-up), INFORMS Computing Society, 2012
arXiv slides
Algorithms: RCDC, UCDC, RCDS
[7] Peter Richtárik and Martin Takáč
Efficiency of randomized coordinate descent methods on minimization problems with a composite objective function
Proceedings of Signal Processing with Adaptive Sparse
Structured Representations, 2011
pdf
Algorithms: UCDC
[6] Peter Richtárik
Finding sparse approximations to extreme eigenvectors: generalized power method for sparse PCA and extensions
Proceedings of
Signal Processing with Adaptive Sparse Structured Representations, 2011
pdf
Algorithms: GPower, ADM
[5] Peter Richtárik
Approximate level method for nonsmooth convex minimization
Journal
of Optimization Theory and Applications 152(2):334–350, 2012
Optimization Online
Algorithms: Approximate Level Method
[4] Michel Journée, Yurii Nesterov, Peter Richtárik and Rodolphe Sepulchre
Generalized power method for sparse principal component analysis
Journal of Machine Learning Research 11:517–553, 2010
arXiv slides poster
Algorithms: GPower
[3] Peter Richtárik
Improved algorithms for convex minimization in relative scale
SIAM Journal on Optimization 21(3):1141–1167, 2011
pdf slides
Algorithms: SubBis, SubSearchNR, SubBisNR, SmoothBis
[2] Peter Richtárik
Simultaneously solving seven optimization problems in relative scale
Technical Report, 2009
Optimization Online
Algorithms: Inc, IncDec
[1] Peter Richtárik
Some algorithms for large-scale convex and linear minimization in relative scale
PhD Dissertation, School of Operations Research and Information Engineering, Cornell University, 2007
Algorithms: SubBis, SmoothBis, Inc, IncDec