Downloads · 30 days
0
PhilippXXY/var-lam
var-lam is a machine learning model from PhilippXXY. Use it for the machine learning task on the model card, and read the license before you ship it in a product.
<h1 align="center"Variable-Array Reconstruction for Latent Acoustic Mapping</h1
Downloads · 30 days
0
Access
Public
Updated Sep 21, 2026
Repo size
18.1 GB
Likes
0
Public
Click a slice to open those files.
.pt17.8 GB · 100%
From the Hugging Face model README
VAR-LAM is a research codebase for Variable-Array Reconstruction for Latent Acoustic Mapping. It predicts complex cross-spectral matrices (CSMs) from multi-channel audio by combining a learned STFT encoder, residual vector quantisation, microphone-pair reasoning, a factorised transformer, and a LAM-style steering-matrix decoder.
<p align="center"> <img src="docs/img/model_architecture.svg" alt="VAR-LAM model architecture" max-width="100%"> </p>Install the source repository and its dependencies:
git clone https://github.com/PhilippXXY/var-lam.git
cd var-lam
uv sync --all-groups
The source repository contains small JSON pointers for published models. Its loader downloads checkpoints from this model repository and reuses the Hugging Face cache.
To download a standalone checkpoint explicitly:
hf download PhilippXXY/var-lam \
checkpoints/32ch-to-32ch/ablations/components/transformer-off.pt \
--revision main \
--local-dir .
Checkpoint paths are organised by channel mapping and experiment:
checkpoints/
├── 32ch-to-32ch/
│ ├── baseline.pt
│ ├── direction-grid/{64,128,256,512,1024,2048,4096,8192}.pt
│ ├── frequency/{linear-0hz-1khz,linear-0hz-4khz,mel-50hz-12khz}.pt
│ ├── input-representation/{magnitude-stft,partial-phat-stft,phat-stft}.pt
│ └── ablations/
│ ├── components/{cross-channel-attention-on,lstm-off-temporal-attention-off,lstm-off-token-comparison-off-transformer-off,token-comparison-no-diagonal,token-comparison-no-differences,token-comparison-no-geometry,token-comparison-no-products,token-comparison-off-transformer-off,token-comparison-off,transformer-off}.pt
│ ├── transformer-attention/{band-only,pair-only}.pt
│ └── rvq/{q2-k1024,q4-k1024,q4-k4096,q8-k64,q8-k256,q8-k512,q8-k2048,q12-k16,q16-k1024}.pt
├── 4ch-to-32ch/
│ ├── baseline.pt
│ ├── direction-grid/{64,128,256,512,1024,2048,4096,8192}.pt
│ ├── input-representation/{magnitude-stft,partial-phat-stft,phat-stft}.pt
│ └── ablations/
│ ├── components/{cross-channel-attention-on,lstm-off-token-comparison-off-transformer-off,rvq-off,token-comparison-off-transformer-off}.pt
│ └── microphone-layout/{horizontal,random,vertical-front}.pt
└── 4ch-to-4ch/
├── baseline.pt
├── input-representation/{magnitude-stft,partial-phat-stft,phat-stft}.pt
└── ablations/components/{cross-channel-attention-on,lstm-off-temporal-attention-off,lstm-off-token-comparison-off-transformer-off,token-comparison-off-transformer-off,token-comparison-off,transformer-off}.pt
transformer-attention/ contains ablations isolated to the factorised transformer's attention axes. Optional modules and combined component ablations remain under components/.
The two direction-grid/128.pt files are aliases of their corresponding baselines, whose decoder direction grid contains 128 points.
Training code, configuration, and instructions live in the source repository. Training writes timestamped latest.pt and best.pt files locally; publish selected outputs here only when they are ready for reuse.
Run a pointer-backed checkpoint through the source repository:
uv run src/infer.py \
--config configs/inference_config.yaml \
--device cuda \
--checkpoint checkpoints/32ch-to-32ch/ablations/components/transformer-off.pt \
--checkpoint-revision main \
--dataset locata
The loader reconstructs the architecture and channel-input mode from checkpoint metadata before loading the state dictionary.
See the rendered VAR-LAM documentation for dataset setup, configuration, training, inference, metrics, and visualisation workflows.