[2024-03-07 13:41:35,149] INFO: Will use torch.nn.parallel.DistributedDataParallel() and 4 gpus [2024-03-07 13:41:35,152] INFO: NVIDIA GeForce GTX 1080 Ti [2024-03-07 13:41:35,152] INFO: NVIDIA GeForce GTX 1080 Ti [2024-03-07 13:41:35,152] INFO: NVIDIA GeForce GTX 1080 Ti [2024-03-07 13:41:35,152] INFO: NVIDIA GeForce GTX 1080 Ti [2024-03-07 13:41:44,885] INFO: using attention_type=efficient [2024-03-07 13:41:44,893] INFO: using attention_type=efficient [2024-03-07 13:41:44,901] INFO: using attention_type=efficient [2024-03-07 13:41:44,909] INFO: using attention_type=efficient [2024-03-07 13:41:44,916] INFO: using attention_type=efficient [2024-03-07 13:41:44,924] INFO: using attention_type=efficient [2024-03-07 13:41:48,510] INFO: DistributedDataParallel( (module): MLPF( (nn0): Sequential( (0): Linear(in_features=42, out_features=256, bias=True) (1): ELU(alpha=1.0) (2): LayerNorm((256,), eps=1e-05, elementwise_affine=True) (3): Dropout(p=0.3, inplace=False) (4): Linear(in_features=256, out_features=256, bias=True) ) (conv_id): ModuleList( (0-2): 3 x SelfAttentionLayer( (mha): MultiheadAttention( (out_proj): NonDynamicallyQuantizableLinear(in_features=256, out_features=256, bias=True) ) (norm0): LayerNorm((256,), eps=1e-05, elementwise_affine=True) (norm1): LayerNorm((256,), eps=1e-05, elementwise_affine=True) (seq): Sequential( (0): Linear(in_features=256, out_features=256, bias=True) (1): ELU(alpha=1.0) (2): Linear(in_features=256, out_features=256, bias=True) (3): ELU(alpha=1.0) ) (dropout): Dropout(p=0.3, inplace=False) ) ) (conv_reg): ModuleList( (0-2): 3 x SelfAttentionLayer( (mha): MultiheadAttention( (out_proj): NonDynamicallyQuantizableLinear(in_features=256, out_features=256, bias=True) ) (norm0): LayerNorm((256,), eps=1e-05, elementwise_affine=True) (norm1): LayerNorm((256,), eps=1e-05, elementwise_affine=True) (seq): Sequential( (0): Linear(in_features=256, out_features=256, bias=True) (1): ELU(alpha=1.0) (2): Linear(in_features=256, out_features=256, bias=True) (3): ELU(alpha=1.0) ) (dropout): Dropout(p=0.3, inplace=False) ) ) (nn_id): Sequential( (0): Linear(in_features=810, out_features=256, bias=True) (1): ELU(alpha=1.0) (2): LayerNorm((256,), eps=1e-05, elementwise_affine=True) (3): Dropout(p=0.3, inplace=False) (4): Linear(in_features=256, out_features=9, bias=True) ) (nn_pt): RegressionOutput( (nn): Sequential( (0): Linear(in_features=819, out_features=256, bias=True) (1): ELU(alpha=1.0) (2): LayerNorm((256,), eps=1e-05, elementwise_affine=True) (3): Dropout(p=0.3, inplace=False) (4): Linear(in_features=256, out_features=2, bias=True) ) ) (nn_eta): RegressionOutput( (nn): Sequential( (0): Linear(in_features=819, out_features=256, bias=True) (1): ELU(alpha=1.0) (2): LayerNorm((256,), eps=1e-05, elementwise_affine=True) (3): Dropout(p=0.3, inplace=False) (4): Linear(in_features=256, out_features=2, bias=True) ) ) (nn_sin_phi): RegressionOutput( (nn): Sequential( (0): Linear(in_features=819, out_features=256, bias=True) (1): ELU(alpha=1.0) (2): LayerNorm((256,), eps=1e-05, elementwise_affine=True) (3): Dropout(p=0.3, inplace=False) (4): Linear(in_features=256, out_features=2, bias=True) ) ) (nn_cos_phi): RegressionOutput( (nn): Sequential( (0): Linear(in_features=819, out_features=256, bias=True) (1): ELU(alpha=1.0) (2): LayerNorm((256,), eps=1e-05, elementwise_affine=True) (3): Dropout(p=0.3, inplace=False) (4): Linear(in_features=256, out_features=2, bias=True) ) ) (nn_energy): RegressionOutput( (nn): Sequential( (0): Linear(in_features=819, out_features=256, bias=True) (1): ELU(alpha=1.0) (2): LayerNorm((256,), eps=1e-05, elementwise_affine=True) (3): Dropout(p=0.3, inplace=False) (4): Linear(in_features=256, out_features=2, bias=True) ) ) (nn_charge): Sequential( (0): Linear(in_features=819, out_features=256, bias=True) (1): ELU(alpha=1.0) (2): LayerNorm((256,), eps=1e-05, elementwise_affine=True) (3): Dropout(p=0.3, inplace=False) (4): Linear(in_features=256, out_features=3, bias=True) ) (nn_probX): Sequential( (0): Linear(in_features=819, out_features=256, bias=True) (1): ELU(alpha=1.0) (2): LayerNorm((256,), eps=1e-05, elementwise_affine=True) (3): Dropout(p=0.3, inplace=False) (4): Linear(in_features=256, out_features=1, bias=True) ) ) ) [2024-03-07 13:41:48,512] INFO: Trainable parameters: 4139031 [2024-03-07 13:41:48,512] INFO: Non-trainable parameters: 0 [2024-03-07 13:41:48,512] INFO: Total parameters: 4139031 [2024-03-07 13:41:48,527] INFO: Modules Trainable params Non-tranable params Trainable Parameters Non-tranable Parameters module.nn0.0.weight NaN NaN 10752.0 - module.nn0.0.bias NaN NaN 256.0 - module.nn0.2.weight NaN NaN 256.0 - module.nn0.2.bias NaN NaN 256.0 - module.nn0.4.weight NaN NaN 65536.0 - module.nn0.4.bias NaN NaN 256.0 - module.conv_id.0.mha.in_proj_weight NaN NaN 196608.0 - module.conv_id.0.mha.in_proj_bias NaN NaN 768.0 - module.conv_id.0.mha.out_proj.weight NaN NaN 65536.0 - module.conv_id.0.mha.out_proj.bias NaN NaN 256.0 - module.conv_id.0.norm0.weight NaN NaN 256.0 - module.conv_id.0.norm0.bias NaN NaN 256.0 - module.conv_id.0.norm1.weight NaN NaN 256.0 - module.conv_id.0.norm1.bias NaN NaN 256.0 - module.conv_id.0.seq.0.weight NaN NaN 65536.0 - module.conv_id.0.seq.0.bias NaN NaN 256.0 - module.conv_id.0.seq.2.weight NaN NaN 65536.0 - module.conv_id.0.seq.2.bias NaN NaN 256.0 - module.conv_id.1.mha.in_proj_weight NaN NaN 196608.0 - module.conv_id.1.mha.in_proj_bias NaN NaN 768.0 - module.conv_id.1.mha.out_proj.weight NaN NaN 65536.0 - module.conv_id.1.mha.out_proj.bias NaN NaN 256.0 - module.conv_id.1.norm0.weight NaN NaN 256.0 - module.conv_id.1.norm0.bias NaN NaN 256.0 - module.conv_id.1.norm1.weight NaN NaN 256.0 - module.conv_id.1.norm1.bias NaN NaN 256.0 - module.conv_id.1.seq.0.weight NaN NaN 65536.0 - module.conv_id.1.seq.0.bias NaN NaN 256.0 - module.conv_id.1.seq.2.weight NaN NaN 65536.0 - module.conv_id.1.seq.2.bias NaN NaN 256.0 - module.conv_id.2.mha.in_proj_weight NaN NaN 196608.0 - module.conv_id.2.mha.in_proj_bias NaN NaN 768.0 - module.conv_id.2.mha.out_proj.weight NaN NaN 65536.0 - module.conv_id.2.mha.out_proj.bias NaN NaN 256.0 - module.conv_id.2.norm0.weight NaN NaN 256.0 - module.conv_id.2.norm0.bias NaN NaN 256.0 - module.conv_id.2.norm1.weight NaN NaN 256.0 - module.conv_id.2.norm1.bias NaN NaN 256.0 - module.conv_id.2.seq.0.weight NaN NaN 65536.0 - module.conv_id.2.seq.0.bias NaN NaN 256.0 - module.conv_id.2.seq.2.weight NaN NaN 65536.0 - module.conv_id.2.seq.2.bias NaN NaN 256.0 - module.conv_reg.0.mha.in_proj_weight NaN NaN 196608.0 - module.conv_reg.0.mha.in_proj_bias NaN NaN 768.0 - module.conv_reg.0.mha.out_proj.weight NaN NaN 65536.0 - module.conv_reg.0.mha.out_proj.bias NaN NaN 256.0 - module.conv_reg.0.norm0.weight NaN NaN 256.0 - module.conv_reg.0.norm0.bias NaN NaN 256.0 - module.conv_reg.0.norm1.weight NaN NaN 256.0 - module.conv_reg.0.norm1.bias NaN NaN 256.0 - module.conv_reg.0.seq.0.weight NaN NaN 65536.0 - module.conv_reg.0.seq.0.bias NaN NaN 256.0 - module.conv_reg.0.seq.2.weight NaN NaN 65536.0 - module.conv_reg.0.seq.2.bias NaN NaN 256.0 - module.conv_reg.1.mha.in_proj_weight NaN NaN 196608.0 - module.conv_reg.1.mha.in_proj_bias NaN NaN 768.0 - module.conv_reg.1.mha.out_proj.weight NaN NaN 65536.0 - module.conv_reg.1.mha.out_proj.bias NaN NaN 256.0 - module.conv_reg.1.norm0.weight NaN NaN 256.0 - module.conv_reg.1.norm0.bias NaN NaN 256.0 - module.conv_reg.1.norm1.weight NaN NaN 256.0 - module.conv_reg.1.norm1.bias NaN NaN 256.0 - module.conv_reg.1.seq.0.weight NaN NaN 65536.0 - module.conv_reg.1.seq.0.bias NaN NaN 256.0 - module.conv_reg.1.seq.2.weight NaN NaN 65536.0 - module.conv_reg.1.seq.2.bias NaN NaN 256.0 - module.conv_reg.2.mha.in_proj_weight NaN NaN 196608.0 - module.conv_reg.2.mha.in_proj_bias NaN NaN 768.0 - module.conv_reg.2.mha.out_proj.weight NaN NaN 65536.0 - module.conv_reg.2.mha.out_proj.bias NaN NaN 256.0 - module.conv_reg.2.norm0.weight NaN NaN 256.0 - module.conv_reg.2.norm0.bias NaN NaN 256.0 - module.conv_reg.2.norm1.weight NaN NaN 256.0 - module.conv_reg.2.norm1.bias NaN NaN 256.0 - module.conv_reg.2.seq.0.weight NaN NaN 65536.0 - module.conv_reg.2.seq.0.bias NaN NaN 256.0 - module.conv_reg.2.seq.2.weight NaN NaN 65536.0 - module.conv_reg.2.seq.2.bias NaN NaN 256.0 - module.nn_id.0.weight NaN NaN 207360.0 - module.nn_id.0.bias NaN NaN 256.0 - module.nn_id.2.weight NaN NaN 256.0 - module.nn_id.2.bias NaN NaN 256.0 - module.nn_id.4.weight NaN NaN 2304.0 - module.nn_id.4.bias NaN NaN 9.0 - module.nn_pt.nn.0.weight NaN NaN 209664.0 - module.nn_pt.nn.0.bias NaN NaN 256.0 - module.nn_pt.nn.2.weight NaN NaN 256.0 - module.nn_pt.nn.2.bias NaN NaN 256.0 - module.nn_pt.nn.4.weight NaN NaN 512.0 - module.nn_pt.nn.4.bias NaN NaN 2.0 - module.nn_eta.nn.0.weight NaN NaN 209664.0 - module.nn_eta.nn.0.bias NaN NaN 256.0 - module.nn_eta.nn.2.weight NaN NaN 256.0 - module.nn_eta.nn.2.bias NaN NaN 256.0 - module.nn_eta.nn.4.weight NaN NaN 512.0 - module.nn_eta.nn.4.bias NaN NaN 2.0 - module.nn_sin_phi.nn.0.weight NaN NaN 209664.0 - module.nn_sin_phi.nn.0.bias NaN NaN 256.0 - module.nn_sin_phi.nn.2.weight NaN NaN 256.0 - module.nn_sin_phi.nn.2.bias NaN NaN 256.0 - module.nn_sin_phi.nn.4.weight NaN NaN 512.0 - module.nn_sin_phi.nn.4.bias NaN NaN 2.0 - module.nn_cos_phi.nn.0.weight NaN NaN 209664.0 - module.nn_cos_phi.nn.0.bias NaN NaN 256.0 - module.nn_cos_phi.nn.2.weight NaN NaN 256.0 - module.nn_cos_phi.nn.2.bias NaN NaN 256.0 - module.nn_cos_phi.nn.4.weight NaN NaN 512.0 - module.nn_cos_phi.nn.4.bias NaN NaN 2.0 - module.nn_energy.nn.0.weight NaN NaN 209664.0 - module.nn_energy.nn.0.bias NaN NaN 256.0 - module.nn_energy.nn.2.weight NaN NaN 256.0 - module.nn_energy.nn.2.bias NaN NaN 256.0 - module.nn_energy.nn.4.weight NaN NaN 512.0 - module.nn_energy.nn.4.bias NaN NaN 2.0 - module.nn_charge.0.weight NaN NaN 209664.0 - module.nn_charge.0.bias NaN NaN 256.0 - module.nn_charge.2.weight NaN NaN 256.0 - module.nn_charge.2.bias NaN NaN 256.0 - module.nn_charge.4.weight NaN NaN 768.0 - module.nn_charge.4.bias NaN NaN 3.0 - module.nn_probX.0.weight NaN NaN 209664.0 - module.nn_probX.0.bias NaN NaN 256.0 - module.nn_probX.2.weight NaN NaN 256.0 - module.nn_probX.2.bias NaN NaN 256.0 - module.nn_probX.4.weight NaN NaN 256.0 - module.nn_probX.4.bias NaN NaN 1.0 - [2024-03-07 13:41:48,551] INFO: Creating experiment dir /pfvol/experiments/MLPF_cms_Transformer_MET_Truepyg-cms-small_20240307_134117_981007 [2024-03-07 13:41:48,551] INFO: Model directory /pfvol/experiments/MLPF_cms_Transformer_MET_Truepyg-cms-small_20240307_134117_981007 [2024-03-07 13:41:49,124] INFO: train_dataset: cms_pf_ttbar, 80000 [2024-03-07 13:41:49,547] INFO: train_dataset: cms_pf_qcd, 80000 [2024-03-07 13:41:49,651] INFO: valid_dataset: cms_pf_ttbar, 20000 [2024-03-07 13:41:49,714] INFO: valid_dataset: cms_pf_qcd, 20000 [2024-03-07 13:41:49,821] INFO: Initiating epoch #1 train run on device rank=0 [2024-03-07 20:29:54,800] INFO: Initiating epoch #1 valid run on device rank=0 [2024-03-07 21:11:29,786] INFO: Rank 0: epoch=1 / 30 train_loss=88.2886 valid_loss=85.1583 stale=0 time=449.67m eta=13040.3m [2024-03-07 21:11:29,951] INFO: Initiating epoch #2 train run on device rank=0 [2024-03-08 03:59:51,301] INFO: Initiating epoch #2 valid run on device rank=0 [2024-03-08 04:47:06,228] INFO: Rank 0: epoch=2 / 30 train_loss=84.5293 valid_loss=84.6377 stale=0 time=455.6m eta=12673.8m [2024-03-08 04:47:07,737] INFO: Initiating epoch #3 train run on device rank=0 [2024-03-08 11:30:09,199] INFO: Initiating epoch #3 valid run on device rank=0 [2024-03-08 12:06:15,604] INFO: Rank 0: epoch=3 / 30 train_loss=84.3273 valid_loss=84.1771 stale=0 time=439.13m eta=12099.9m [2024-03-08 12:06:16,777] INFO: Initiating epoch #4 train run on device rank=0 [2024-03-08 18:51:08,256] INFO: Initiating epoch #4 valid run on device rank=0 [2024-03-08 19:31:46,958] INFO: Rank 0: epoch=4 / 30 train_loss=84.1490 valid_loss=84.1387 stale=0 time=445.5m eta=11634.7m [2024-03-08 19:31:48,051] INFO: Initiating epoch #5 train run on device rank=0 [2024-03-09 02:28:52,862] INFO: Initiating epoch #5 valid run on device rank=0 [2024-03-09 03:04:59,811] INFO: Rank 0: epoch=5 / 30 train_loss=84.0877 valid_loss=84.0735 stale=0 time=453.2m eta=11215.8m [2024-03-09 03:05:01,181] INFO: Initiating epoch #6 train run on device rank=0 [2024-03-09 09:50:19,469] INFO: Initiating epoch #6 valid run on device rank=0 [2024-03-09 10:31:36,489] INFO: Rank 0: epoch=6 / 30 train_loss=83.9862 valid_loss=84.1129 stale=1 time=446.59m eta=10759.1m [2024-03-09 10:31:37,407] INFO: Initiating epoch #7 train run on device rank=0 [2024-03-09 17:16:57,475] INFO: Initiating epoch #7 valid run on device rank=0 [2024-03-09 18:05:41,424] INFO: Rank 0: epoch=7 / 30 train_loss=83.9237 valid_loss=83.9716 stale=0 time=454.07m eta=10329.8m [2024-03-09 18:05:43,160] INFO: Initiating epoch #8 train run on device rank=0 [2024-03-10 00:59:26,082] INFO: Initiating epoch #8 valid run on device rank=0 [2024-03-10 01:50:52,495] INFO: Rank 0: epoch=8 / 30 train_loss=83.8370 valid_loss=83.9389 stale=0 time=465.16m eta=9924.9m [2024-03-10 01:50:53,808] INFO: Initiating epoch #9 train run on device rank=0