BEGIN:VCALENDAR
VERSION:2.0
PRODID:icalendar-ruby
CALSCALE:GREGORIAN
METHOD:PUBLISH
BEGIN:VTIMEZONE
TZID:Europe/Vienna
BEGIN:DAYLIGHT
DTSTART:20260329T030000
TZOFFSETFROM:+0100
TZOFFSETTO:+0200
RRULE:FREQ=YEARLY;BYDAY=-1SU;BYMONTH=3
TZNAME:CEST
END:DAYLIGHT
BEGIN:STANDARD
DTSTART:20261025T020000
TZOFFSETFROM:+0200
TZOFFSETTO:+0100
RRULE:FREQ=YEARLY;BYDAY=-1SU;BYMONTH=10
TZNAME:CET
END:STANDARD
END:VTIMEZONE
BEGIN:VEVENT
DTSTAMP:20260803T135916Z
UID:1786975200@ist.ac.at
DTSTART:20260817T160000
DTEND:20260817T170000
DESCRIPTION:Speaker: Kevin Kögler\nhosted by Krzysztof Pietrzak\nAbstract:
  Modern datasets often contain as many variables as observations\, or many
  more\, making classical statistical tools difficult to apply and standard
  machine-learning models hard to understand. To address this problem\, thi
 s thesis develops methods for inference and compression in this high-dimen
 sional regime\, centered on approximate message passing (AMP)\, a family o
 f iterative inference algorithms which are conjecture to be optimal for ma
 ny of these problems.First\, we introduce rotationally invariant generaliz
 ed AMP (RI-GAMP)\, extending AMP beyond independent Gaussian design matric
 es to a broad class of matrices with complex correlations and arbitrary li
 miting spectra. We characterize its performance rigorously through state e
 volution and show that it approaches the performance of Vector AMP while a
 voiding its computationally expensive singular-value decomposition.We then
  study nonlinear autoencoders when the input and compressed dimensions gro
 w proportionally. For Gaussian data\, we derive fundamental limits for one
 -bit compression by shallow two-layer autoencoders\, characterize their op
 timal solutions\, and prove that gradient methods reach them. For structur
 ed data\, however\, a shallow linear decoder may ignore information such a
 s sparsity and behave exactly as it does for Gaussian inputs. Adding nonli
 nearities or depth overcomes this limitation: deeper decoders can learn co
 mputations closely related to AMP and approach optimal reconstruction for 
 a fixed encoder.Finally\, we apply these ideas to predict plasma-protein a
 bundance from UK Biobank individual-level genetic data\, using approximate
 ly 8.4 million genetic variants from roughly 40\,000 individuals. A two-st
 age genetic Vector AMP pipeline first selects a small set of predictive ge
 nomic regions and then fits more detailed linear or tree-based models. Com
 bined with a one-hot representation of the genotype data\, this procedure 
 closes the previously reported average gap between linear and deep-learnin
 g prediction and improves prediction for many proteins in a proteome-wide 
 comparison. 
LOCATION:Office Bldg West / Ground floor / Heinzel Seminar Room (I21.EG.101
 ) and Zoom\, ISTA
ORGANIZER:
SUMMARY:Kevin Kögler: Thesis Defense: Compression and Inference in High Di
 mensions: An Approximate Message Passing Perspective
URL:https://talks-calendar.ista.ac.at/events/6577
END:VEVENT
END:VCALENDAR
