icefall

Author	SHA1	Message	Date
Daniel Povey	76e66408c5	Some cosmetic improvements	2022-09-27 11:08:44 +08:00
Daniel Povey	e25ca74955	Use a measure of correlation for eigs that can be negative.	2022-07-26 13:40:57 +08:00
Daniel Povey	b9696878b4	Update diagnostics stats	2022-07-26 12:39:51 +08:00
Daniel Povey	7e88e2a0e9	Increase debug freq; add type to diagnostics and increase precision of mean,rms	2022-07-17 06:40:16 +08:00
Daniel Povey	ca09b9798f	Remove decomposition code from checkpoint.py; restore double precision model_avg	2022-06-01 14:01:58 +08:00
Daniel Povey	da2ffd4d27	Do average computation in double precision	2022-05-31 14:39:21 +08:00
Daniel Povey	b2259184b5	Use single precision for model average; increase average-period to 200.	2022-05-31 14:31:46 +08:00
Daniel Povey	8d4c987e21	Update checkpoint.py to support decompose argument	2022-05-31 14:25:45 +08:00
Daniel Povey	7011956c6c	Merge remote-tracking branch 'upstream/master' into cain3d_clean_merge	2022-05-31 12:17:45 +08:00
LIyong.Guo	c4ee2bc0af	[Ready to merge]stateless6: states4 + hubert distillation. (#387 ) * a copy of stateless4 as base * distillation with hubert * fix typo * example usage * usage * Update egs/librispeech/ASR/pruned_transducer_stateless6/hubert_xlarge.py Co-authored-by: Fangjun Kuang <csukuangfj@gmail.com> * fix comment * add results of 100hours * Update egs/librispeech/ASR/pruned_transducer_stateless6/hubert_xlarge.py Co-authored-by: Fangjun Kuang <csukuangfj@gmail.com> * Update egs/librispeech/ASR/pruned_transducer_stateless6/hubert_xlarge.py Co-authored-by: Fangjun Kuang <csukuangfj@gmail.com> * check fairseq and quantization * a short intro to distillation framework * Update egs/librispeech/ASR/pruned_transducer_stateless6/hubert_xlarge.py Co-authored-by: Fangjun Kuang <csukuangfj@gmail.com> * add intro of statless6 in README * fix type error of dst_manifest_dir * Update egs/librispeech/ASR/pruned_transducer_stateless6/hubert_xlarge.py Co-authored-by: Fangjun Kuang <csukuangfj@gmail.com> * make export.py call stateless6/train.py instead of stateless2/train.py * update results by stateless6 * adjust results format * fix typo Co-authored-by: Fangjun Kuang <csukuangfj@gmail.com>	2022-05-28 12:37:50 +08:00
Daniel Povey	8e454bcf9e	Exclude size=500 dim from projection; try to use double for model average	2022-05-26 15:15:27 +08:00
Mingshuang Luo	ec5a112831	[Ready to merge] Do some coding style checks for the latest files (#379 ) * style check * do changes for .flake8 * a change for compute_fbank_yesno.py	2022-05-20 19:30:38 +08:00
Daniel Povey	5230e73e41	Small fixes	2022-05-19 12:49:00 +08:00
Daniel Povey	c0fdfabaf3	Remove memory-limit options arg	2022-05-19 11:30:56 +08:00
Daniel Povey	c2c46ea023	Update diagnostics, hopefully print more stats. # Conflicts: # egs/librispeech/ASR/pruned_transducer_stateless4b/train.py	2022-05-19 11:29:31 +08:00
Fangjun Kuang	cd460f7bf1	Stringify torch.__version__ before serializing it. (#354 )	2022-05-07 17:18:34 +08:00
Zengwei Yao	20f092e709	Support decoding with averaged model when using --iter (#353 ) * support decoding with averaged model when using --iter * minor fix * monir fix of copyright date	2022-05-07 13:09:11 +08:00
Zengwei Yao	c059ef3169	Keep model_avg on cpu (#348 ) * keep model_avg on cpu * explicitly convert model_avg to cpu * minor fix * remove device convertion for model_avg * modify usage of the model device in train.py * change model.device to next(model.parameters()).device for decoding * assert params.start_epoch>0 * assert params.start_epoch>0, params.start_epoch	2022-05-07 10:42:34 +08:00
Zengwei Yao	00c48ec1f3	Model average (#344 ) * First upload of model average codes. * minor fix * update decode file * update .flake8 * rename pruned_transducer_stateless3 to pruned_transducer_stateless4 * change epoch number counter starting from 1 instead of 0 * minor fix of pruned_transducer_stateless4/train.py * refactor the checkpoint.py * minor fix, update docs, and modify the epoch number to count from 1 in the pruned_transducer_stateless4/decode.py * update author info * add docs of the scaling in function average_checkpoints_with_averaged_model	2022-05-05 21:20:04 +08:00
Fangjun Kuang	9aeea3e1af	Support averaging models with weight tying. (#333 )	2022-04-26 13:32:03 +08:00
Wang, Guanbo	5fe58de43c	GigaSpeech recipe (#120 ) * initial commit * support download, data prep, and fbank * on-the-fly feature extraction by default * support BPE based lang * support HLG for BPE * small fix * small fix * chunked feature extraction by default * Compute features for GigaSpeech by splitting the manifest. * Fixes after review. * Split manifests into 2000 pieces. * set audio duration mismatch tolerance to 0.01 * small fix * add conformer training recipe * Add conformer.py without pre-commit checking * lazy loading and use SingleCutSampler * DynamicBucketingSampler * use KaldifeatFbank to compute fbank for musan * use pretrained language model and lexicon * use 3gram to decode, 4gram to rescore * Add decode.py * Update .flake8 * Delete compute_fbank_gigaspeech.py * Use BucketingSampler for valid and test dataloader * Update params in train.py * Use bpe_500 * update params in decode.py * Decrease num_paths while CUDA OOM * Added README * Update RESULTS * black * Decrease num_paths while CUDA OOM * Decode with post-processing * Update results * Remove lazy_load option * Use default `storage_type` * Keep the original tolerance * Use split-lazy * black * Update pretrained model Co-authored-by: Fangjun Kuang <csukuangfj@gmail.com>	2022-04-14 16:07:22 +08:00
Guo Liyong	78418ac37c	fix comments	2022-04-13 13:09:24 +08:00
Mingshuang Luo	93c60a9d30	Code style check for librispeech pruned transducer stateless2 (#308 )	2022-04-11 22:15:18 +08:00
Daniel Povey	6eb6d9b4cd	Merge pull request #288 from danpovey/reworked_model Reworked model	2022-04-11 15:03:08 +08:00
Wei Kang	f721a2fd7a	Minor fixes for logging (#296 ) * Minor fixes for logging * Minor fix	2022-04-10 23:34:18 +08:00
Zengwei Yao	08473a17aa	Modify init (#301 ) * update icefall/__init__.py to import more common functions. * update icefall/__init__.py * make imports style consistent. * exclude black check for icefall/__init__.py in pyproject.toml.	2022-04-10 23:29:28 +08:00
Daniel Povey	d1e4ae788d	Refactor how learning rate is set.	2022-04-10 15:25:27 +08:00
Fangjun Kuang	7c0070e6f6	Display torch version in the training log. (#299 )	2022-04-08 11:39:54 +08:00
Zengwei Yao	ceeb95bcb8	update icefall/__init__.py to import more common functions. (#294 )	2022-04-06 11:55:29 +08:00
Fangjun Kuang	87cf9231ea	Support specifying iteration number of checkpoints for decoding. (#289 )	2022-04-03 13:02:08 +08:00
Zengwei Yao	0b6a2213c3	Modify icefall/__init__.py. (#287 ) * Modify icefall/__init__.py to import common functions defined in icefall/utils.py. * Modify icefall/__init__.py and .flake8.	2022-04-02 15:01:45 +08:00
LIyong.Guo	fc40bfea82	fix typo of torch.eig (#281 ) Co-authored-by: glynpu <glynwpu@qq.com>	2022-03-31 10:43:46 +08:00
Mingshuang Luo	f686635b54	Update diagnostics (#260 ) * update diagnostics.py	2022-03-30 14:52:55 +08:00
Fangjun Kuang	ae564f91e6	Periodically saving checkpoint after processing given number of batches (#259 ) * Periodically saving checkpoint after processing given number of batches.	2022-03-20 23:51:33 +08:00
Mingshuang Luo	518ec6414a	Update diagnostics.py (#254 ) * update diagnostics.py * do some changes	2022-03-16 20:17:45 +08:00
yaozengwei	ad62981765	Add diagnostics (#230 ) * Adding diagnostics code... * Move diagnostics code from local dir to the shared icefall dir * Remove the diagnostics code in the local dir * Update docs of arguments, and remove stats_types() function in TensorDiagnosticOptions object. * Update docs of arguments. * Add copyright information. * Corrected the time in copyright information. Co-authored-by: Daniel Povey <dpovey@gmail.com>	2022-03-04 15:38:23 +08:00
Fangjun Kuang	cbf8c18ebd	Minor fixes for aishell (#218 ) * Minor fixes to aishell. * Minor fixes.	2022-02-19 22:28:19 +08:00
Wei Kang	b702281e90	Use k2 pruned transducer loss to train conformer-transducer model (#194 ) * Using k2 pruned version transducer loss to train model * Fix style * Minor fixes	2022-02-17 13:33:54 +08:00
Wang, Guanbo	e8eb408760	Incremental pruning threshold (#214 ) * Incremental pruning threshold * flake8 * black * minor fix	2022-02-16 16:59:27 +08:00
Wang, Guanbo	be1c86b06c	print num_frame as %.2f (#204 )	2022-02-08 14:56:58 +08:00
Piotr Żelasko	f92c24a73a	Merge branch 'master' into feature/libri-conformer-phone-ctc	2022-01-24 10:18:56 -05:00
Piotr Żelasko	f0f35e6671	black	2022-01-21 17:22:41 -05:00
Piotr Żelasko	3d109b121d	Remove train_phones.py and modify train.py instead	2022-01-21 17:08:53 -05:00
huangruizhe	298faabb90	minor fixes	2022-01-02 23:38:33 -08:00
huangruizhe	7577b08bed	fixed the mistake	2022-01-02 23:32:43 -08:00
huangruizhe	82c8fac6ee	fixed a case where BOW can have problem to compute (ZeroDivisionError)	2022-01-02 15:29:50 -08:00
huangruizhe	0a67015d63	Update make_kn_lm.py	2022-01-02 00:27:27 -08:00
huangruizhe	49aab7e658	Update make_kn_lm.py Fixed issue #163	2022-01-02 00:14:27 -08:00
Fangjun Kuang	95af039733	RNN-T training for yesno. (#141 ) * RNN-T training for yesno. * Rename Jointer to Joiner.	2021-12-07 21:44:37 +08:00
Fangjun Kuang	ec591698b0	Associate a cut with token alignment (without repeats) (#125 ) * WIP: Associate a cut with token alignment (without repeats) * Save framewise alignments with/without repeats. * Minor fixes.	2021-11-29 18:50:54 +08:00

1 2

87 Commits