mirror of
https://github.com/SakanaAI/doc-to-lora.git
synced 2026-07-23 17:01:04 +02:00
eval longbench cmd
This commit is contained in:
parent
dde5f5cb4b
commit
47edd5df1e
1 changed files with 7 additions and 0 deletions
|
|
@ -40,4 +40,11 @@ run python intx_sft.py configs/math_and_code.yaml --model_name_or_path=google/ge
|
|||
### HyperLoRA finetune on GSM8K
|
||||
```bash
|
||||
run python intx_sft.py configs/gsm8k.yaml --model_name_or_path=google/gemma-2-2b-it --num_train_epochs=5 --per_device_train_batch_size=16 --gradient_accumulation_steps=1 --per_device_eval_batch_size=8 --exp_setup=hyper_lora --aggregator_type=perceiver --target_modules=down_proj --num_blocks=1 --num_self_attends_per_block=16 --self_attention_widening_factor=1 --eval_steps=5000 --save_steps=5000 --learning_rate=2e-5 --neftune_noise_alpha=5 --use_light_weight_lora=True --light_weight_latent_size=512 --load_best_model_at_end=True --metric_for_best_model=pwc_loss --add_negative_prompt=False --add_repeat_prompt=False --ctx_encoder_model_name_or_path=meta-llama/Llama-3.2-11B-Vision-Instruct
|
||||
```
|
||||
|
||||
### Evaluation
|
||||
LongBench
|
||||
```bash
|
||||
cd LongBench/LongBench
|
||||
run python eval_ctx_to_lora.py --model_name Mar16_12-38-01_slurm0-a3nodeset-12_54818_32426662/checkpoint-136782 --checkpoint_path ../../train_outputs/runs/Mar16_12-38-01_slurm0-a3nodeset-12_54818_32426662/checkpoint-136782/pytorch_model.bin
|
||||
```
|
||||
Loading…
Add table
Add a link
Reference in a new issue