Daniel Kleine 
							
						 
					 
					
						
						
						
						
							
						
						
							479b0e2aa9 
							
						 
					 
					
						
						
							
							fixed gqa qkv code comments ( #660 )  
						
						
						
						
					 
					
						2025-06-13 08:21:28 -05:00 
						 
				 
			
				
					
						
							
							
								Pratyush Subhadarshi 
							
						 
					 
					
						
						
						
						
							
						
						
							d142741ef4 
							
						 
					 
					
						
						
							
							Correcting the wrong reference ( #649 )  
						
						... 
						
						
						
						Co-authored-by: Sebastian Raschka <mail@sebastianraschka.com> 
						
						
					 
					
						2025-06-12 16:35:51 -05:00 
						 
				 
			
				
					
						
							
							
								Sebastian Raschka 
							
						 
					 
					
						
						
						
						
							
						
						
							a3c4c33347 
							
						 
					 
					
						
						
							
							Reduce Llama 3 RoPE memory requirements ( #658 )  
						
						... 
						
						
						
						* Llama3 from scratch improvements
* Fix Llama 3 expensive RoPE memory issue
* updates
* update package
* benchmark
* remove unused rescale_theta 
						
						
					 
					
						2025-06-12 11:08:02 -05:00 
						 
				 
			
				
					
						
							
							
								Sebastian Raschka 
							
						 
					 
					
						
						
						
						
							
						
						
							3eca919a52 
							
						 
					 
					
						
						
							
							Llama3 from scratch improvements ( #621 )  
						
						... 
						
						
						
						* Llama3 from scratch improvements
* restore 
						
						
					 
					
						2025-04-16 18:08:26 -05:00 
						 
				 
			
				
					
						
							
							
								Henry Shi 
							
						 
					 
					
						
						
						
						
							
						
						
							88250d953d 
							
						 
					 
					
						
						
							
							updated exercise 5.3 ( #615 )  
						
						... 
						
						
						
						* updated exercise 5.3
temperature can be set to 0 to regardless of top_k setting to force deterministic behavior
* fix notebook json
---------
Co-authored-by: Sebastian Raschka <mail@sebastianraschka.com> 
						
						
					 
					
						2025-04-13 13:06:57 -05:00 
						 
				 
			
				
					
						
							
							
								Sebastian Raschka 
							
						 
					 
					
						
						
						
						
							
						
						
							97a199e40b 
							
						 
					 
					
						
						
							
							Disable mask saving as weight in Llama 3 model ( #604 )  
						
						... 
						
						
						
						* Disable mask saving as weight
* update pixi
* update pixi 
						
						
					 
					
						2025-04-06 09:33:36 -05:00 
						 
				 
			
				
					
						
							
							
								Sebastian Raschka 
							
						 
					 
					
						
						
						
						
							
						
						
							c43d7ef663 
							
						 
					 
					
						
						
							
							reformat nbs ( #602 )  
						
						
						
						
					 
					
						2025-04-05 16:18:27 -05:00 
						 
				 
			
				
					
						
							
							
								Sebastian Raschka 
							
						 
					 
					
						
						
						
						
							
						
						
							396e96ab07 
							
						 
					 
					
						
						
							
							Fix Llama language typo in bonus materials ( #597 )  
						
						
						
						
					 
					
						2025-04-02 21:41:36 -05:00 
						 
				 
			
				
					
						
							
							
								Sebastian Raschka 
							
						 
					 
					
						
						
						
						
							
						
						
							4128a91c1d 
							
						 
					 
					
						
						
							
							Add Llama 3.2 to pkg ( #591 )  
						
						... 
						
						
						
						* Add Llama 3.2 to pkg
* remove redundant attributes
* update tests
* updates
* updates
* updates
* fix link
* fix link 
						
						
					 
					
						2025-03-31 18:59:47 -05:00 
						 
				 
			
				
					
						
							
							
								casinca 
							
						 
					 
					
						
						
						
						
							
						
						
							d7c316533a 
							
						 
					 
					
						
						
							
							removing unused RoPE parameters ( #590 )  
						
						... 
						
						
						
						* removing unused RoPE parameters
* remove redundant context_length in GQA
---------
Co-authored-by: Sebastian Raschka <mail@sebastianraschka.com> 
						
						
					 
					
						2025-03-31 17:10:39 -05:00 
						 
				 
			
				
					
						
							
							
								Sebastian Raschka 
							
						 
					 
					
						
						
						
						
							
						
						
							4e3b752e5e 
							
						 
					 
					
						
						
							
							Memory optimized Llama ( #588 )  
						
						... 
						
						
						
						* Memory optimized Llama
* re-ad login 
						
						
					 
					
						2025-03-30 15:18:12 -05:00 
						 
				 
			
				
					
						
							
							
								Sebastian Raschka 
							
						 
					 
					
						
						
						
						
							
						
						
							e55e3e88e1 
							
						 
					 
					
						
						
							
							Alt weight loading code via PyTorch ( #585 )  
						
						... 
						
						
						
						* Alt weight loading code via PyTorch
* commit additional files 
						
						
					 
					
						2025-03-27 20:10:23 -05:00 
						 
				 
			
				
					
						
							
							
								Sebastian Raschka 
							
						 
					 
					
						
						
						
						
							
						
						
							c9271ac427 
							
						 
					 
					
						
						
							
							Adjust comment to save compiled model ( #583 )  
						
						
						
						
					 
					
						2025-03-27 10:43:45 -05:00 
						 
				 
			
				
					
						
							
							
								Sebastian Raschka 
							
						 
					 
					
						
						
						
						
							
						
						
							857acfcc12 
							
						 
					 
					
						
						
							
							Vocab padding clarification ( #582 )  
						
						... 
						
						
						
						* vocab padding clarification
* Update ch05/10_llm-training-speed/README.md 
						
						
					 
					
						2025-03-26 13:19:55 -05:00 
						 
				 
			
				
					
						
							
							
								Sebastian Raschka 
							
						 
					 
					
						
						
						
						
							
						
						
							fee7d4bb05 
							
						 
					 
					
						
						
							
							More explicit torchrun usage doc ( #578 )  
						
						
						
						
					 
					
						2025-03-24 12:01:03 -05:00 
						 
				 
			
				
					
						
							
							
								Sebastian Raschka 
							
						 
					 
					
						
						
						
						
							
						
						
							cf6fb73553 
							
						 
					 
					
						
						
							
							Add readme ( #577 )  
						
						
						
						
					 
					
						2025-03-23 19:35:12 -05:00 
						 
				 
			
				
					
						
							
							
								Sebastian Raschka 
							
						 
					 
					
						
						
						
						
							
						
						
							7114ccd10d 
							
						 
					 
					
						
						
							
							Add PyPI package ( #576 )  
						
						... 
						
						
						
						* Add PyPI package
* fixes
* fixes 
						
						
					 
					
						2025-03-23 19:28:49 -05:00 
						 
				 
			
				
					
						
							
							
								Sebastian Raschka 
							
						 
					 
					
						
						
						
						
							
						
						
							85f2bc0a58 
							
						 
					 
					
						
						
							
							Speed comparison figure ( #575 )  
						
						
						
						
					 
					
						2025-03-21 11:29:49 -05:00 
						 
				 
			
				
					
						
							
							
								Greg Gandenberger 
							
						 
					 
					
						
						
						
						
							
						
						
							1ec5631c70 
							
						 
					 
					
						
						
							
							Fix minor printing issue and note inconsistency across platforms ( #563 )  
						
						... 
						
						
						
						* Fix printing issue and note inconsistency
* Rerun notebook 
						
						
					 
					
						2025-03-14 15:12:09 -05:00 
						 
				 
			
				
					
						
							
							
								Sebastian Raschka 
							
						 
					 
					
						
						
						
						
							
						
						
							4fb0ea9d1f 
							
						 
					 
					
						
						
							
							Specify UTF-8 encoding in the json load command explicitely ( #557 )  
						
						
						
						
					 
					
						2025-03-05 11:46:21 -06:00 
						 
				 
			
				
					
						
							
							
								Sebastian Raschka 
							
						 
					 
					
						
						
						
						
							
						
						
							de60da9a6b 
							
						 
					 
					
						
						
							
							Add a note about "zsh: illegal hardware instruction python" error ( #555 )  
						
						
						
						
					 
					
						2025-03-02 15:18:24 -06:00 
						 
				 
			
				
					
						
							
							
								Sebastian Raschka 
							
						 
					 
					
						
						
						
						
							
						
						
							fa5760a8de 
							
						 
					 
					
						
						
							
							GitHub markdown updates ( #545 )  
						
						... 
						
						
						
						* GitHub markdown updates
* Apply suggestions from code review
* Apply suggestions from code review 
						
						
					 
					
						2025-02-23 12:25:44 -06:00 
						 
				 
			
				
					
						
							
							
								Sebastian Raschka 
							
						 
					 
					
						
						
						
						
							
						
						
							5016499d1d 
							
						 
					 
					
						
						
							
							Uv workflow improvements ( #531 )  
						
						... 
						
						
						
						* Uv workflow improvements
* Uv workflow improvements
* linter improvements
* pytproject.toml fixes
* pytproject.toml fixes
* pytproject.toml fixes
* pytproject.toml fixes
* pytproject.toml fixes
* pytproject.toml fixes
* windows fixes
* windows fixes
* windows fixes
* windows fixes
* windows fixes
* windows fixes
* win32 fix
* win32 fix
* win32 fix
* win32 fix
* win32 fix
* win32 fix
* win32 fix
* win32 fix
* win32 fix
* win32 fix
* win32 fix
* win32 fix
* win32 fix
* win32 fix
* win32 fix
* win32 fix
* win32 fix
* win32 fix
* win32 fix 
						
						
					 
					
						2025-02-16 13:16:51 -06:00 
						 
				 
			
				
					
						
							
							
								Sebastian Raschka 
							
						 
					 
					
						
						
						
						
							
						
						
							e818be42e1 
							
						 
					 
					
						
						
							
							Update link to vocab size increase ( #526 )  
						
						... 
						
						
						
						* Update link to vocab size increase
* Update ch05/10_llm-training-speed/README.md
* Update ch05/10_llm-training-speed/README.md 
						
						
					 
					
						2025-02-14 08:03:01 -06:00 
						 
				 
			
				
					
						
							
							
								Sebastian Raschka 
							
						 
					 
					
						
						
						
						
							
						
						
							6370898ce6 
							
						 
					 
					
						
						
							
							PyTorch tips for better training performance ( #525 )  
						
						... 
						
						
						
						* PyTorch tips for better training performance
* formatting
* pep 8 
						
						
					 
					
						2025-02-12 16:10:34 -06:00 
						 
				 
			
				
					
						
							
							
								Sebastian Raschka 
							
						 
					 
					
						
						
						
						
							
						
						
							9dce43ec31 
							
						 
					 
					
						
						
							
							Upgrade to NumPy 2.0 ( #520 )  
						
						... 
						
						
						
						* Upgrade to NumPy 2.0
* bump pytorch
* bump pytorch
* bump pytorch
* bump pytorch
* bump pytorch
* update
* update packages 
						
						
					 
					
						2025-02-09 06:21:58 -06:00 
						 
				 
			
				
					
						
							
							
								Sebastian Raschka 
							
						 
					 
					
						
						
						
						
							
						
						
							5efa731c0f 
							
						 
					 
					
						
						
							
							Mention small discrepancy due to Dropout non-reproducibility in PyTorch ( #519 )  
						
						... 
						
						
						
						* Mention small discrepancy due to Dropout non-reproducibility in PyTorch
* bump pytorch version 
						
						
					 
					
						2025-02-06 14:59:52 -06:00 
						 
				 
			
				
					
						
							
							
								Sebastian Raschka 
							
						 
					 
					
						
						
						
						
							
						
						
							fd24a3679a 
							
						 
					 
					
						
						
							
							Alternative weight loading via .safetensors ( #507 )  
						
						
						
						
					 
					
						2025-01-29 08:15:29 -06:00 
						 
				 
			
				
					
						
							
							
								Sebastian Raschka 
							
						 
					 
					
						
						
						
						
							
						
						
							dcaac28b92 
							
						 
					 
					
						
						
							
							Bonus material: extending tokenizers ( #496 )  
						
						... 
						
						
						
						* Bonus material: extending tokenizers
* small wording update 
						
						
					 
					
						2025-01-22 09:26:54 -06:00 
						 
				 
			
				
					
						
							
							
								Sebastian Raschka 
							
						 
					 
					
						
						
						
						
							
						
						
							992f3068d1 
							
						 
					 
					
						
						
							
							Auto download DPO dataset if not already available in path ( #479 )  
						
						... 
						
						
						
						* Auto download DPO dataset if not already available in path
* update tests to account for latest HF transformers release in unit tests
* pep 8 
						
						
					 
					
						2025-01-12 12:27:28 -06:00 
						 
				 
			
				
					
						
							
							
								Sebastian Raschka 
							
						 
					 
					
						
						
						
						
							
						
						
							7659af7cdd 
							
						 
					 
					
						
						
							
							Add backup URL for gpt2 weights ( #469 )  
						
						... 
						
						
						
						* Add backup URL for gpt2 weights
* newline 
						
						
					 
					
						2025-01-05 11:28:09 -06:00 
						 
				 
			
				
					
						
							
							
								Sebastian Raschka 
							
						 
					 
					
						
						
						
						
							
						
						
							05a816e270 
							
						 
					 
					
						
						
							
							fix misplaced parenthesis and update license ( #466 )  
						
						
						
						
					 
					
						2025-01-04 11:14:08 -06:00 
						 
				 
			
				
					
						
							
							
								casinca 
							
						 
					 
					
						
						
						
						
							
						
						
							57fdd94358 
							
						 
					 
					
						
						
							
							[minor] typo & comments ( #441 )  
						
						... 
						
						
						
						* typo & comment
- safe -> save
- commenting code: batch_size, seq_len = in_idx.shape
* comment
- adding # NEW for assert num_heads % num_kv_groups == 0
* update memory wording
---------
Co-authored-by: rasbt <mail@sebastianraschka.com> 
						
						
					 
					
						2024-11-18 19:52:42 +09:00 
						 
				 
			
				
					
						
							
							
								Sebastian Raschka 
							
						 
					 
					
						
						
						
						
							
						
						
							129d0d740f 
							
						 
					 
					
						
						
							
							Add missing device transfer in gpt_generate.py ( #436 )  
						
						
						
						
					 
					
						2024-11-14 19:12:53 +09:00 
						 
				 
			
				
					
						
							
							
								Daniel Kleine 
							
						 
					 
					
						
						
						
						
							
						
						
							7e6f8ce020 
							
						 
					 
					
						
						
							
							updated RoPE statement ( #423 )  
						
						... 
						
						
						
						* updated RoPE statement
* updated .gitignore
* Update ch05/07_gpt_to_llama/converting-gpt-to-llama2.ipynb
---------
Co-authored-by: Sebastian Raschka <mail@sebastianraschka.com> 
						
						
					 
					
						2024-10-30 08:00:08 -05:00 
						 
				 
			
				
					
						
							
							
								ROHAN WINSOR 
							
						 
					 
					
						
						
						
						
							
						
						
							e85d154522 
							
						 
					 
					
						
						
							
							Fix argument name in LlamaTokenizer constructor ( #421 )  
						
						... 
						
						
						
						This PR addresses an oversight in the LlamaTokenizer class by changing the constructor argument from filepath to tokenizer_file. 
						
						
					 
					
						2024-10-29 18:01:36 -05:00 
						 
				 
			
				
					
						
							
							
								Daniel Kleine 
							
						 
					 
					
						
						
						
						
							
						
						
							2b24a7ef30 
							
						 
					 
					
						
						
							
							minor fixes: Llama 3.2 standalone ( #420 )  
						
						... 
						
						
						
						* minor fixes
* reformat rope base as float
---------
Co-authored-by: rasbt <mail@sebastianraschka.com> 
						
						
					 
					
						2024-10-25 21:08:06 -05:00 
						 
				 
			
				
					
						
							
							
								Sebastian Raschka 
							
						 
					 
					
						
						
						
						
							
						
						
							75ede3e340 
							
						 
					 
					
						
						
							
							RoPE theta rescaling ( #419 )  
						
						... 
						
						
						
						* rope fixes
* update
* update
* cleanup 
						
						
					 
					
						2024-10-25 15:27:23 -05:00 
						 
				 
			
				
					
						
							
							
								Daniel Kleine 
							
						 
					 
					
						
						
						
						
							
						
						
							0ed1e0d099 
							
						 
					 
					
						
						
							
							fixed typos ( #414 )  
						
						... 
						
						
						
						* fixed typos
* fixed formatting
* Update ch03/02_bonus_efficient-multihead-attention/mha-implementations.ipynb
* del weights after load into model
---------
Co-authored-by: Sebastian Raschka <mail@sebastianraschka.com> 
						
						
					 
					
						2024-10-24 18:23:53 -05:00 
						 
				 
			
				
					
						
							
							
								Daniel Kleine 
							
						 
					 
					
						
						
						
						
							
						
						
							8b60460319 
							
						 
					 
					
						
						
							
							Updated Llama 2 to 3 paths ( #413 )  
						
						... 
						
						
						
						* llama 2 and 3 path fixes
* updated llama 3, 3.1 and 3.2 paths
* updated .gitignore
* Typo fix
---------
Co-authored-by: Sebastian Raschka <mail@sebastianraschka.com> 
						
						
					 
					
						2024-10-24 07:40:08 -05:00 
						 
				 
			
				
					
						
							
							
								Sebastian Raschka 
							
						 
					 
					
						
						
						
						
							
						
						
							632d7772b2 
							
						 
					 
					
						
						
							
							Update test-requirements-extra.txt  
						
						
						
						
					 
					
						2024-10-23 19:19:58 -05:00 
						 
				 
			
				
					
						
							
							
								Sebastian Raschka 
							
						 
					 
					
						
						
						
						
							
						
						
							f8bdfe12e1 
							
						 
					 
					
						
						
							
							RoPE updates ( #412 )  
						
						... 
						
						
						
						* RoPE updates
* Apply suggestions from code review
* updates
* updates
* updates 
						
						
					 
					
						2024-10-23 18:07:49 -05:00 
						 
				 
			
				
					
						
							
							
								Sebastian Raschka 
							
						 
					 
					
						
						
						
						
							
						
						
							6dd3fbd79d 
							
						 
					 
					
						
						
							
							Update tests.py  
						
						
						
						
					 
					
						2024-10-23 07:48:33 -05:00 
						 
				 
			
				
					
						
							
							
								Sebastian Raschka 
							
						 
					 
					
						
						
						
						
							
						
						
							9726ca6546 
							
						 
					 
					
						
						
							
							RoPE increase ( #407 )  
						
						
						
						
					 
					
						2024-10-21 19:58:38 -05:00 
						 
				 
			
				
					
						
							
							
								rasbt 
							
						 
					 
					
						
						
						
						
							
						
						
							3567fb656d 
							
						 
					 
					
						
						
							
							update mmap section  
						
						
						
						
					 
					
						2024-10-14 14:27:19 -05:00 
						 
				 
			
				
					
						
							
							
								rasbt 
							
						 
					 
					
						
						
						
						
							
						
						
							31fb74133a 
							
						 
					 
					
						
						
							
							add mmap=True comparison  
						
						
						
						
					 
					
						2024-10-14 11:09:55 -05:00 
						 
				 
			
				
					
						
							
							
								Sebastian Raschka 
							
						 
					 
					
						
						
						
						
							
						
						
							3d54af20f5 
							
						 
					 
					
						
						
							
							Memory efficient weight loading ( #401 )  
						
						... 
						
						
						
						* memory efficient weight loading
* remove unused code 
						
						
					 
					
						2024-10-14 10:30:25 -05:00 
						 
				 
			
				
					
						
							
							
								Sebastian Raschka 
							
						 
					 
					
						
						
						
						
							
						
						
							6a9bedc2ec 
							
						 
					 
					
						
						
							
							Update bonus section formatting ( #400 )  
						
						
						
						
					 
					
						2024-10-12 10:26:08 -05:00 
						 
				 
			
				
					
						
							
							
								Sebastian Raschka 
							
						 
					 
					
						
						
						
						
							
						
						
							37db3f0913 
							
						 
					 
					
						
						
							
							Add Llama 3.2 RoPE to CI ( #391 )  
						
						... 
						
						
						
						* add Llama 3.2 RoPE to CI
* update 
						
						
					 
					
						2024-10-08 08:28:34 -05:00 
						 
				 
			
				
					
						
							
							
								Sebastian Raschka 
							
						 
					 
					
						
						
						
						
							
						
						
							06604f4b84 
							
						 
					 
					
						
						
							
							Introduce buffers to improve Llama 3.2 efficiency ( #389 )  
						
						... 
						
						
						
						* Introduce buffers to improve Llama 3.2 efficiency
* update
* update 
						
						
					 
					
						2024-10-06 12:49:04 -05:00