LLMs-from-scratch

mirror of https://github.com/rasbt/LLMs-from-scratch.git synced 2025-11-27 23:52:23 +00:00

Author	SHA1	Message	Date
Sebastian Raschka	984cca3f64	Fix code comment: embed_dim -> d_out (#698 )	2025-06-22 16:36:39 -05:00
Sebastian Raschka	a5ea296259	Use more recent sentencepiece tokenizer API (#696 )	2025-06-22 13:52:30 -05:00
Sebastian Raschka	2351a1f282	Fix some wording issues in the notes (#695 )	2025-06-22 13:46:16 -05:00
Sebastian Raschka	0b15a00574	Qwen3 KV cache (#688 )	2025-06-21 17:34:39 -05:00
Sebastian Raschka	9d62ca0598	Llama 3 KV Cache (#685 ) * Llama 3 KV Cache * skip expensive tests on Gh actions * Update __init__.py	2025-06-21 10:55:20 -05:00
Sebastian Raschka	9f6f514191	Fix formatting in Qwen3 nb (#680 ) * Fix formatting in Qwen3 nb * upd	2025-06-20 07:28:27 -05:00
Daniel Kleine	e79bb50a1b	fixed plot_losses (#677 )	2025-06-19 18:55:43 -05:00
Sebastian Raschka	3d4bce6d57	Qwen3 From Scratch (#678 ) * Qwen3 From Scratch * rev other file * upd * upd * upd * url fixes	2025-06-19 18:44:38 -05:00
casinca	e700c66b7a	removed old args in GQA class (#674 )	2025-06-17 13:09:53 -05:00
Daniel Kleine	479b0e2aa9	fixed gqa qkv code comments (#660 )	2025-06-13 08:21:28 -05:00
Pratyush Subhadarshi	d142741ef4	Correcting the wrong reference (#649 ) Co-authored-by: Sebastian Raschka <mail@sebastianraschka.com>	2025-06-12 16:35:51 -05:00
Sebastian Raschka	a3c4c33347	Reduce Llama 3 RoPE memory requirements (#658 ) * Llama3 from scratch improvements * Fix Llama 3 expensive RoPE memory issue * updates * update package * benchmark * remove unused rescale_theta	2025-06-12 11:08:02 -05:00
Sebastian Raschka	3eca919a52	Llama3 from scratch improvements (#621 ) * Llama3 from scratch improvements * restore	2025-04-16 18:08:26 -05:00
Henry Shi	88250d953d	updated exercise 5.3 (#615 ) * updated exercise 5.3 temperature can be set to 0 to regardless of top_k setting to force deterministic behavior * fix notebook json --------- Co-authored-by: Sebastian Raschka <mail@sebastianraschka.com>	2025-04-13 13:06:57 -05:00
Sebastian Raschka	97a199e40b	Disable mask saving as weight in Llama 3 model (#604 ) * Disable mask saving as weight * update pixi * update pixi	2025-04-06 09:33:36 -05:00
Sebastian Raschka	c43d7ef663	reformat nbs (#602 )	2025-04-05 16:18:27 -05:00
Sebastian Raschka	396e96ab07	Fix Llama language typo in bonus materials (#597 )	2025-04-02 21:41:36 -05:00
Sebastian Raschka	4128a91c1d	Add Llama 3.2 to pkg (#591 ) * Add Llama 3.2 to pkg * remove redundant attributes * update tests * updates * updates * updates * fix link * fix link	2025-03-31 18:59:47 -05:00
casinca	d7c316533a	removing unused RoPE parameters (#590 ) * removing unused RoPE parameters * remove redundant context_length in GQA --------- Co-authored-by: Sebastian Raschka <mail@sebastianraschka.com>	2025-03-31 17:10:39 -05:00
Sebastian Raschka	4e3b752e5e	Memory optimized Llama (#588 ) * Memory optimized Llama * re-ad login	2025-03-30 15:18:12 -05:00
Sebastian Raschka	e55e3e88e1	Alt weight loading code via PyTorch (#585 ) * Alt weight loading code via PyTorch * commit additional files	2025-03-27 20:10:23 -05:00
Sebastian Raschka	c9271ac427	Adjust comment to save compiled model (#583 )	2025-03-27 10:43:45 -05:00
Sebastian Raschka	857acfcc12	Vocab padding clarification (#582 ) * vocab padding clarification * Update ch05/10_llm-training-speed/README.md	2025-03-26 13:19:55 -05:00
Sebastian Raschka	fee7d4bb05	More explicit torchrun usage doc (#578 )	2025-03-24 12:01:03 -05:00
Sebastian Raschka	cf6fb73553	Add readme (#577 )	2025-03-23 19:35:12 -05:00
Sebastian Raschka	7114ccd10d	Add PyPI package (#576 ) * Add PyPI package * fixes * fixes	2025-03-23 19:28:49 -05:00
Sebastian Raschka	85f2bc0a58	Speed comparison figure (#575 )	2025-03-21 11:29:49 -05:00
Greg Gandenberger	1ec5631c70	Fix minor printing issue and note inconsistency across platforms (#563 ) * Fix printing issue and note inconsistency * Rerun notebook	2025-03-14 15:12:09 -05:00
Sebastian Raschka	4fb0ea9d1f	Specify UTF-8 encoding in the json load command explicitely (#557 )	2025-03-05 11:46:21 -06:00
Sebastian Raschka	de60da9a6b	Add a note about "zsh: illegal hardware instruction python" error (#555 )	2025-03-02 15:18:24 -06:00
Sebastian Raschka	fa5760a8de	GitHub markdown updates (#545 ) * GitHub markdown updates * Apply suggestions from code review * Apply suggestions from code review	2025-02-23 12:25:44 -06:00
Sebastian Raschka	5016499d1d	Uv workflow improvements (#531 ) * Uv workflow improvements * Uv workflow improvements * linter improvements * pytproject.toml fixes * pytproject.toml fixes * pytproject.toml fixes * pytproject.toml fixes * pytproject.toml fixes * pytproject.toml fixes * windows fixes * windows fixes * windows fixes * windows fixes * windows fixes * windows fixes * win32 fix * win32 fix * win32 fix * win32 fix * win32 fix * win32 fix * win32 fix * win32 fix * win32 fix * win32 fix * win32 fix * win32 fix * win32 fix * win32 fix * win32 fix * win32 fix * win32 fix * win32 fix * win32 fix	2025-02-16 13:16:51 -06:00
Sebastian Raschka	e818be42e1	Update link to vocab size increase (#526 ) * Update link to vocab size increase * Update ch05/10_llm-training-speed/README.md * Update ch05/10_llm-training-speed/README.md	2025-02-14 08:03:01 -06:00
Sebastian Raschka	6370898ce6	PyTorch tips for better training performance (#525 ) * PyTorch tips for better training performance * formatting * pep 8	2025-02-12 16:10:34 -06:00
Sebastian Raschka	9dce43ec31	Upgrade to NumPy 2.0 (#520 ) * Upgrade to NumPy 2.0 * bump pytorch * bump pytorch * bump pytorch * bump pytorch * bump pytorch * update * update packages	2025-02-09 06:21:58 -06:00
Sebastian Raschka	5efa731c0f	Mention small discrepancy due to Dropout non-reproducibility in PyTorch (#519 ) * Mention small discrepancy due to Dropout non-reproducibility in PyTorch * bump pytorch version	2025-02-06 14:59:52 -06:00
Sebastian Raschka	fd24a3679a	Alternative weight loading via .safetensors (#507 )	2025-01-29 08:15:29 -06:00
Sebastian Raschka	dcaac28b92	Bonus material: extending tokenizers (#496 ) * Bonus material: extending tokenizers * small wording update	2025-01-22 09:26:54 -06:00
Sebastian Raschka	992f3068d1	Auto download DPO dataset if not already available in path (#479 ) * Auto download DPO dataset if not already available in path * update tests to account for latest HF transformers release in unit tests * pep 8	2025-01-12 12:27:28 -06:00
Sebastian Raschka	7659af7cdd	Add backup URL for gpt2 weights (#469 ) * Add backup URL for gpt2 weights * newline	2025-01-05 11:28:09 -06:00
Sebastian Raschka	05a816e270	fix misplaced parenthesis and update license (#466 )	2025-01-04 11:14:08 -06:00
casinca	57fdd94358	[minor] typo & comments (#441 ) * typo & comment - safe -> save - commenting code: batch_size, seq_len = in_idx.shape * comment - adding # NEW for assert num_heads % num_kv_groups == 0 * update memory wording --------- Co-authored-by: rasbt <mail@sebastianraschka.com>	2024-11-18 19:52:42 +09:00
Sebastian Raschka	129d0d740f	Add missing device transfer in gpt_generate.py (#436 )	2024-11-14 19:12:53 +09:00
Daniel Kleine	7e6f8ce020	updated RoPE statement (#423 ) * updated RoPE statement * updated .gitignore * Update ch05/07_gpt_to_llama/converting-gpt-to-llama2.ipynb --------- Co-authored-by: Sebastian Raschka <mail@sebastianraschka.com>	2024-10-30 08:00:08 -05:00
ROHAN WINSOR	e85d154522	Fix argument name in LlamaTokenizer constructor (#421 ) This PR addresses an oversight in the LlamaTokenizer class by changing the constructor argument from filepath to tokenizer_file.	2024-10-29 18:01:36 -05:00
Daniel Kleine	2b24a7ef30	minor fixes: Llama 3.2 standalone (#420 ) * minor fixes * reformat rope base as float --------- Co-authored-by: rasbt <mail@sebastianraschka.com>	2024-10-25 21:08:06 -05:00
Sebastian Raschka	75ede3e340	RoPE theta rescaling (#419 ) * rope fixes * update * update * cleanup	2024-10-25 15:27:23 -05:00
Daniel Kleine	0ed1e0d099	fixed typos (#414 ) * fixed typos * fixed formatting * Update ch03/02_bonus_efficient-multihead-attention/mha-implementations.ipynb * del weights after load into model --------- Co-authored-by: Sebastian Raschka <mail@sebastianraschka.com>	2024-10-24 18:23:53 -05:00
Daniel Kleine	8b60460319	Updated Llama 2 to 3 paths (#413 ) * llama 2 and 3 path fixes * updated llama 3, 3.1 and 3.2 paths * updated .gitignore * Typo fix --------- Co-authored-by: Sebastian Raschka <mail@sebastianraschka.com>	2024-10-24 07:40:08 -05:00
Sebastian Raschka	632d7772b2	Update test-requirements-extra.txt	2024-10-23 19:19:58 -05:00

1 2 3 4

186 Commits