Sebastian Raschka 
							
						 
					 
					
						
						
						
						
							
						
						
							85f2bc0a58 
							
						 
					 
					
						
						
							
							Speed comparison figure ( #575 )  
						
						
						
						
					 
					
						2025-03-21 11:29:49 -05:00 
						 
				 
			
				
					
						
							
							
								Greg Gandenberger 
							
						 
					 
					
						
						
						
						
							
						
						
							1ec5631c70 
							
						 
					 
					
						
						
							
							Fix minor printing issue and note inconsistency across platforms ( #563 )  
						
						... 
						
						
						
						* Fix printing issue and note inconsistency
* Rerun notebook 
						
						
					 
					
						2025-03-14 15:12:09 -05:00 
						 
				 
			
				
					
						
							
							
								Sebastian Raschka 
							
						 
					 
					
						
						
						
						
							
						
						
							4fb0ea9d1f 
							
						 
					 
					
						
						
							
							Specify UTF-8 encoding in the json load command explicitely ( #557 )  
						
						
						
						
					 
					
						2025-03-05 11:46:21 -06:00 
						 
				 
			
				
					
						
							
							
								Sebastian Raschka 
							
						 
					 
					
						
						
						
						
							
						
						
							de60da9a6b 
							
						 
					 
					
						
						
							
							Add a note about "zsh: illegal hardware instruction python" error ( #555 )  
						
						
						
						
					 
					
						2025-03-02 15:18:24 -06:00 
						 
				 
			
				
					
						
							
							
								Sebastian Raschka 
							
						 
					 
					
						
						
						
						
							
						
						
							fa5760a8de 
							
						 
					 
					
						
						
							
							GitHub markdown updates ( #545 )  
						
						... 
						
						
						
						* GitHub markdown updates
* Apply suggestions from code review
* Apply suggestions from code review 
						
						
					 
					
						2025-02-23 12:25:44 -06:00 
						 
				 
			
				
					
						
							
							
								Sebastian Raschka 
							
						 
					 
					
						
						
						
						
							
						
						
							5016499d1d 
							
						 
					 
					
						
						
							
							Uv workflow improvements ( #531 )  
						
						... 
						
						
						
						* Uv workflow improvements
* Uv workflow improvements
* linter improvements
* pytproject.toml fixes
* pytproject.toml fixes
* pytproject.toml fixes
* pytproject.toml fixes
* pytproject.toml fixes
* pytproject.toml fixes
* windows fixes
* windows fixes
* windows fixes
* windows fixes
* windows fixes
* windows fixes
* win32 fix
* win32 fix
* win32 fix
* win32 fix
* win32 fix
* win32 fix
* win32 fix
* win32 fix
* win32 fix
* win32 fix
* win32 fix
* win32 fix
* win32 fix
* win32 fix
* win32 fix
* win32 fix
* win32 fix
* win32 fix
* win32 fix 
						
						
					 
					
						2025-02-16 13:16:51 -06:00 
						 
				 
			
				
					
						
							
							
								Sebastian Raschka 
							
						 
					 
					
						
						
						
						
							
						
						
							e818be42e1 
							
						 
					 
					
						
						
							
							Update link to vocab size increase ( #526 )  
						
						... 
						
						
						
						* Update link to vocab size increase
* Update ch05/10_llm-training-speed/README.md
* Update ch05/10_llm-training-speed/README.md 
						
						
					 
					
						2025-02-14 08:03:01 -06:00 
						 
				 
			
				
					
						
							
							
								Sebastian Raschka 
							
						 
					 
					
						
						
						
						
							
						
						
							6370898ce6 
							
						 
					 
					
						
						
							
							PyTorch tips for better training performance ( #525 )  
						
						... 
						
						
						
						* PyTorch tips for better training performance
* formatting
* pep 8 
						
						
					 
					
						2025-02-12 16:10:34 -06:00 
						 
				 
			
				
					
						
							
							
								Sebastian Raschka 
							
						 
					 
					
						
						
						
						
							
						
						
							9dce43ec31 
							
						 
					 
					
						
						
							
							Upgrade to NumPy 2.0 ( #520 )  
						
						... 
						
						
						
						* Upgrade to NumPy 2.0
* bump pytorch
* bump pytorch
* bump pytorch
* bump pytorch
* bump pytorch
* update
* update packages 
						
						
					 
					
						2025-02-09 06:21:58 -06:00 
						 
				 
			
				
					
						
							
							
								Sebastian Raschka 
							
						 
					 
					
						
						
						
						
							
						
						
							5efa731c0f 
							
						 
					 
					
						
						
							
							Mention small discrepancy due to Dropout non-reproducibility in PyTorch ( #519 )  
						
						... 
						
						
						
						* Mention small discrepancy due to Dropout non-reproducibility in PyTorch
* bump pytorch version 
						
						
					 
					
						2025-02-06 14:59:52 -06:00 
						 
				 
			
				
					
						
							
							
								Sebastian Raschka 
							
						 
					 
					
						
						
						
						
							
						
						
							fd24a3679a 
							
						 
					 
					
						
						
							
							Alternative weight loading via .safetensors ( #507 )  
						
						
						
						
					 
					
						2025-01-29 08:15:29 -06:00 
						 
				 
			
				
					
						
							
							
								Sebastian Raschka 
							
						 
					 
					
						
						
						
						
							
						
						
							dcaac28b92 
							
						 
					 
					
						
						
							
							Bonus material: extending tokenizers ( #496 )  
						
						... 
						
						
						
						* Bonus material: extending tokenizers
* small wording update 
						
						
					 
					
						2025-01-22 09:26:54 -06:00 
						 
				 
			
				
					
						
							
							
								Sebastian Raschka 
							
						 
					 
					
						
						
						
						
							
						
						
							992f3068d1 
							
						 
					 
					
						
						
							
							Auto download DPO dataset if not already available in path ( #479 )  
						
						... 
						
						
						
						* Auto download DPO dataset if not already available in path
* update tests to account for latest HF transformers release in unit tests
* pep 8 
						
						
					 
					
						2025-01-12 12:27:28 -06:00 
						 
				 
			
				
					
						
							
							
								Sebastian Raschka 
							
						 
					 
					
						
						
						
						
							
						
						
							7659af7cdd 
							
						 
					 
					
						
						
							
							Add backup URL for gpt2 weights ( #469 )  
						
						... 
						
						
						
						* Add backup URL for gpt2 weights
* newline 
						
						
					 
					
						2025-01-05 11:28:09 -06:00 
						 
				 
			
				
					
						
							
							
								Sebastian Raschka 
							
						 
					 
					
						
						
						
						
							
						
						
							05a816e270 
							
						 
					 
					
						
						
							
							fix misplaced parenthesis and update license ( #466 )  
						
						
						
						
					 
					
						2025-01-04 11:14:08 -06:00 
						 
				 
			
				
					
						
							
							
								casinca 
							
						 
					 
					
						
						
						
						
							
						
						
							57fdd94358 
							
						 
					 
					
						
						
							
							[minor] typo & comments ( #441 )  
						
						... 
						
						
						
						* typo & comment
- safe -> save
- commenting code: batch_size, seq_len = in_idx.shape
* comment
- adding # NEW for assert num_heads % num_kv_groups == 0
* update memory wording
---------
Co-authored-by: rasbt <mail@sebastianraschka.com> 
						
						
					 
					
						2024-11-18 19:52:42 +09:00 
						 
				 
			
				
					
						
							
							
								Sebastian Raschka 
							
						 
					 
					
						
						
						
						
							
						
						
							129d0d740f 
							
						 
					 
					
						
						
							
							Add missing device transfer in gpt_generate.py ( #436 )  
						
						
						
						
					 
					
						2024-11-14 19:12:53 +09:00 
						 
				 
			
				
					
						
							
							
								Daniel Kleine 
							
						 
					 
					
						
						
						
						
							
						
						
							7e6f8ce020 
							
						 
					 
					
						
						
							
							updated RoPE statement ( #423 )  
						
						... 
						
						
						
						* updated RoPE statement
* updated .gitignore
* Update ch05/07_gpt_to_llama/converting-gpt-to-llama2.ipynb
---------
Co-authored-by: Sebastian Raschka <mail@sebastianraschka.com> 
						
						
					 
					
						2024-10-30 08:00:08 -05:00 
						 
				 
			
				
					
						
							
							
								ROHAN WINSOR 
							
						 
					 
					
						
						
						
						
							
						
						
							e85d154522 
							
						 
					 
					
						
						
							
							Fix argument name in LlamaTokenizer constructor ( #421 )  
						
						... 
						
						
						
						This PR addresses an oversight in the LlamaTokenizer class by changing the constructor argument from filepath to tokenizer_file. 
						
						
					 
					
						2024-10-29 18:01:36 -05:00 
						 
				 
			
				
					
						
							
							
								Daniel Kleine 
							
						 
					 
					
						
						
						
						
							
						
						
							2b24a7ef30 
							
						 
					 
					
						
						
							
							minor fixes: Llama 3.2 standalone ( #420 )  
						
						... 
						
						
						
						* minor fixes
* reformat rope base as float
---------
Co-authored-by: rasbt <mail@sebastianraschka.com> 
						
						
					 
					
						2024-10-25 21:08:06 -05:00 
						 
				 
			
				
					
						
							
							
								Sebastian Raschka 
							
						 
					 
					
						
						
						
						
							
						
						
							75ede3e340 
							
						 
					 
					
						
						
							
							RoPE theta rescaling ( #419 )  
						
						... 
						
						
						
						* rope fixes
* update
* update
* cleanup 
						
						
					 
					
						2024-10-25 15:27:23 -05:00 
						 
				 
			
				
					
						
							
							
								Daniel Kleine 
							
						 
					 
					
						
						
						
						
							
						
						
							0ed1e0d099 
							
						 
					 
					
						
						
							
							fixed typos ( #414 )  
						
						... 
						
						
						
						* fixed typos
* fixed formatting
* Update ch03/02_bonus_efficient-multihead-attention/mha-implementations.ipynb
* del weights after load into model
---------
Co-authored-by: Sebastian Raschka <mail@sebastianraschka.com> 
						
						
					 
					
						2024-10-24 18:23:53 -05:00 
						 
				 
			
				
					
						
							
							
								Daniel Kleine 
							
						 
					 
					
						
						
						
						
							
						
						
							8b60460319 
							
						 
					 
					
						
						
							
							Updated Llama 2 to 3 paths ( #413 )  
						
						... 
						
						
						
						* llama 2 and 3 path fixes
* updated llama 3, 3.1 and 3.2 paths
* updated .gitignore
* Typo fix
---------
Co-authored-by: Sebastian Raschka <mail@sebastianraschka.com> 
						
						
					 
					
						2024-10-24 07:40:08 -05:00 
						 
				 
			
				
					
						
							
							
								Sebastian Raschka 
							
						 
					 
					
						
						
						
						
							
						
						
							632d7772b2 
							
						 
					 
					
						
						
							
							Update test-requirements-extra.txt  
						
						
						
						
					 
					
						2024-10-23 19:19:58 -05:00 
						 
				 
			
				
					
						
							
							
								Sebastian Raschka 
							
						 
					 
					
						
						
						
						
							
						
						
							f8bdfe12e1 
							
						 
					 
					
						
						
							
							RoPE updates ( #412 )  
						
						... 
						
						
						
						* RoPE updates
* Apply suggestions from code review
* updates
* updates
* updates 
						
						
					 
					
						2024-10-23 18:07:49 -05:00 
						 
				 
			
				
					
						
							
							
								Sebastian Raschka 
							
						 
					 
					
						
						
						
						
							
						
						
							6dd3fbd79d 
							
						 
					 
					
						
						
							
							Update tests.py  
						
						
						
						
					 
					
						2024-10-23 07:48:33 -05:00 
						 
				 
			
				
					
						
							
							
								Sebastian Raschka 
							
						 
					 
					
						
						
						
						
							
						
						
							9726ca6546 
							
						 
					 
					
						
						
							
							RoPE increase ( #407 )  
						
						
						
						
					 
					
						2024-10-21 19:58:38 -05:00 
						 
				 
			
				
					
						
							
							
								rasbt 
							
						 
					 
					
						
						
						
						
							
						
						
							3567fb656d 
							
						 
					 
					
						
						
							
							update mmap section  
						
						
						
						
					 
					
						2024-10-14 14:27:19 -05:00 
						 
				 
			
				
					
						
							
							
								rasbt 
							
						 
					 
					
						
						
						
						
							
						
						
							31fb74133a 
							
						 
					 
					
						
						
							
							add mmap=True comparison  
						
						
						
						
					 
					
						2024-10-14 11:09:55 -05:00 
						 
				 
			
				
					
						
							
							
								Sebastian Raschka 
							
						 
					 
					
						
						
						
						
							
						
						
							3d54af20f5 
							
						 
					 
					
						
						
							
							Memory efficient weight loading ( #401 )  
						
						... 
						
						
						
						* memory efficient weight loading
* remove unused code 
						
						
					 
					
						2024-10-14 10:30:25 -05:00 
						 
				 
			
				
					
						
							
							
								Sebastian Raschka 
							
						 
					 
					
						
						
						
						
							
						
						
							6a9bedc2ec 
							
						 
					 
					
						
						
							
							Update bonus section formatting ( #400 )  
						
						
						
						
					 
					
						2024-10-12 10:26:08 -05:00 
						 
				 
			
				
					
						
							
							
								Sebastian Raschka 
							
						 
					 
					
						
						
						
						
							
						
						
							37db3f0913 
							
						 
					 
					
						
						
							
							Add Llama 3.2 RoPE to CI ( #391 )  
						
						... 
						
						
						
						* add Llama 3.2 RoPE to CI
* update 
						
						
					 
					
						2024-10-08 08:28:34 -05:00 
						 
				 
			
				
					
						
							
							
								Sebastian Raschka 
							
						 
					 
					
						
						
						
						
							
						
						
							06604f4b84 
							
						 
					 
					
						
						
							
							Introduce buffers to improve Llama 3.2 efficiency ( #389 )  
						
						... 
						
						
						
						* Introduce buffers to improve Llama 3.2 efficiency
* update
* update 
						
						
					 
					
						2024-10-06 12:49:04 -05:00 
						 
				 
			
				
					
						
							
							
								Daniel Kleine 
							
						 
					 
					
						
						
						
						
							
						
						
							4f9775d91c 
							
						 
					 
					
						
						
							
							fixed Llama 2 to 3.2 NBs ( #388 )  
						
						... 
						
						
						
						* updated requirements
* fixes llama2 to llama3
* fixed llama 3.2 standalone
* fixed typo
* fixed rope formula
* Update requirements-extra.txt
* Update ch05/07_gpt_to_llama/converting-llama2-to-llama3.ipynb
* Update ch05/07_gpt_to_llama/converting-llama2-to-llama3.ipynb
* Update ch05/07_gpt_to_llama/standalone-llama32.ipynb
---------
Co-authored-by: Sebastian Raschka <mail@sebastianraschka.com> 
						
						
					 
					
						2024-10-06 09:56:55 -05:00 
						 
				 
			
				
					
						
							
							
								Sebastian Raschka 
							
						 
					 
					
						
						
						
						
							
						
						
							81053ccadd 
							
						 
					 
					
						
						
							
							Add a note about weight tying in Llama 3.2 ( #386 )  
						
						
						
						
					 
					
						2024-10-05 09:20:54 -05:00 
						 
				 
			
				
					
						
							
							
								Sebastian Raschka 
							
						 
					 
					
						
						
						
						
							
						
						
							58c3bb3d9d 
							
						 
					 
					
						
						
							
							Llama 3 ( #384 )  
						
						... 
						
						
						
						* Implement Llama 3.2
* Add Llama 3.2 files
* exclude IMDB link because stanford website seems down 
						
						
					 
					
						2024-10-05 07:52:15 -05:00 
						 
				 
			
				
					
						
							
							
								Sebastian Raschka 
							
						 
					 
					
						
						
						
						
							
						
						
							8d6b25785d 
							
						 
					 
					
						
						
							
							Llama 3.2 requirements file  
						
						
						
						
					 
					
						2024-10-05 07:32:43 -05:00 
						 
				 
			
				
					
						
							
							
								Sebastian Raschka 
							
						 
					 
					
						
						
						
						
							
						
						
							6f86c78763 
							
						 
					 
					
						
						
							
							Implement Llama 3.2 ( #383 )  
						
						
						
						
					 
					
						2024-10-05 07:30:47 -05:00 
						 
				 
			
				
					
						
							
							
								Sebastian Raschka 
							
						 
					 
					
						
						
						
						
							
						
						
							d313f61c86 
							
						 
					 
					
						
						
							
							Cos-sin fix in Llama 2 bonus notebook ( #381 )  
						
						
						
						
					 
					
						2024-10-03 20:45:40 -05:00 
						 
				 
			
				
					
						
							
							
								Sebastian Raschka 
							
						 
					 
					
						
						
						
						
							
						
						
							feb0647c79 
							
						 
					 
					
						
						
							
							Improve rope settings for llama3 ( #380 )  
						
						
						
						
					 
					
						2024-10-03 08:29:54 -05:00 
						 
				 
			
				
					
						
							
							
								rasbt 
							
						 
					 
					
						
						
						
						
							
						
						
							2ae4ad15ba 
							
						 
					 
					
						
						
							
							add section numbers  
						
						
						
						
					 
					
						2024-09-30 08:42:22 -05:00 
						 
				 
			
				
					
						
							
							
								rasbt 
							
						 
					 
					
						
						
						
						
							
						
						
							58d0ce83a4 
							
						 
					 
					
						
						
							
							llama note  
						
						
						
						
					 
					
						2024-09-26 07:41:11 -05:00 
						 
				 
			
				
					
						
							
							
								Sebastian Raschka 
							
						 
					 
					
						
						
						
						
							
						
						
							b8497c1bf5 
							
						 
					 
					
						
						
							
							Add llama2 unit tests ( #372 )  
						
						... 
						
						
						
						* add llama2 unit tests
* update
* updates
* updates
* update file path
* update requirements file
* rmsnorm test
* update 
						
						
					 
					
						2024-09-25 19:40:36 -05:00 
						 
				 
			
				
					
						
							
							
								rasbt 
							
						 
					 
					
						
						
						
						
							
						
						
							a23fca84d5 
							
						 
					 
					
						
						
							
							improve formatting  
						
						
						
						
					 
					
						2024-09-24 18:49:17 -05:00 
						 
				 
			
				
					
						
							
							
								Daniel Kleine 
							
						 
					 
					
						
						
						
						
							
						
						
							4541177063 
							
						 
					 
					
						
						
							
							ch05/07 gpt_to_llama text improvements ( #369 )  
						
						... 
						
						
						
						* fixed typo
* fixed RMSnorm formula
* fixed SwiGLU formula
* temperature=0 for untrained model for reproducibility
* added extra info hf token 
						
						
					 
					
						2024-09-24 18:45:49 -05:00 
						 
				 
			
				
					
						
							
							
								rasbt 
							
						 
					 
					
						
						
						
						
							
						
						
							941629d2c7 
							
						 
					 
					
						
						
							
							add json import  
						
						
						
						
					 
					
						2024-09-23 09:12:35 -05:00 
						 
				 
			
				
					
						
							
							
								rasbt 
							
						 
					 
					
						
						
						
						
							
						
						
							835832a0f9 
							
						 
					 
					
						
						
							
							move access token to config.json  
						
						
						
						
					 
					
						2024-09-23 08:56:16 -05:00 
						 
				 
			
				
					
						
							
							
								rasbt 
							
						 
					 
					
						
						
						
						
							
						
						
							5e6c7230ac 
							
						 
					 
					
						
						
							
							add llama3 comparison  
						
						
						
						
					 
					
						2024-09-23 08:17:10 -05:00 
						 
				 
			
				
					
						
							
							
								Sebastian Raschka 
							
						 
					 
					
						
						
						
						
							
						
						
							c38b003aa9 
							
						 
					 
					
						
						
							
							GPT to Llama ( #368 )  
						
						... 
						
						
						
						* GPT to Llama
* fix urls 
						
						
					 
					
						2024-09-23 07:34:06 -05:00 
						 
				 
			
				
					
						
							
							
								Sebastian Raschka 
							
						 
					 
					
						
						
						
						
							
						
						
							7a9a17608d 
							
						 
					 
					
						
						
							
							Add user interface to ch06 and ch07 ( #366 )  
						
						... 
						
						
						
						* Add user interface to ch06 and ch07
* pep8
* fix url 
						
						
					 
					
						2024-09-21 20:33:00 -05:00