LightRAG

mirror of https://github.com/HKUDS/LightRAG.git synced 2025-11-22 21:15:52 +00:00

Author	SHA1	Message	Date
yangdx	fc8ca1a706	Fix: add muti-process lock for initialize and drop method for all storage	2025-08-12 04:25:09 +08:00
yangdx	44204abef7	Fix linting	2025-08-10 10:59:32 +08:00
yangdx	eb2320e556	Fix: Initialize first_stage_tasks and entity_relation_task to prevent empty-task cancel errors - Initialize first_stage_tasks = [] and entity_relation_task = None at coroutine start - Ensure cancel block safely handles no-op when tasks lists are empty	2025-08-10 10:45:41 +08:00
yangdx	cf064579ce	Remove deprecated keyword extraction query methods - Delete query_with_keywords function - Remove kg_query_with_keywords helper - Drop separate keyword extraction methods	2025-08-08 14:59:39 +08:00
yangdx	c22315ea6d	refactor: remove selective LLM cache clearing functionality - Remove optional 'modes' parameter from aclear_cache() and clear_cache() methods - Replace deprecated drop_cache_by_modes() with drop() method for complete cache clearing - Update API endpoint to ignore mode-specific parameters and clear all cache - Simplify frontend clearCache() function to send empty request body This change ensures all LLM cache is cleared together.	2025-08-05 23:51:51 +08:00
yangdx	01bce8c26e	feat: add warning logs for deleting non-completed documents	2025-08-05 12:21:08 +08:00
yangdx	63496698a1	Fix: ensure data migration is handled by single-process - Wrap migration logic with get_data_init_lock() to ensure single-process execution - Prevent race conditions when multiple processes start simultaneously	2025-08-04 01:47:20 +08:00
yangdx	bf9a6d699b	Fix(lightrag): Handle undirected edges in data migration The `_migrate_entity_relation_data` function previously processed directed edges from `get_all_edges`, which could lead to duplicates (e.g., (A,B) and (B,A)) and an incorrect relation count. This commit normalizes edges by sorting their source and target nodes before adding them to the relation set. This ensures all edges are treated as undirected and are properly deduplicated.	2025-08-03 22:14:24 +08:00
yangdx	e8d8afa846	Removed auto storage management from LightRAG instance creation - The `initialize_storages` method must be explicitly called after LightRAG creation. The `finalize_storages` method should be called before LightRAG destyoyed. - Added explicit data migration check	2025-08-03 12:42:57 +08:00
yangdx	06efab4af2	Revert "Remove auto_manage_storages_states option" This reverts commit bfe6657b316f7e50bc9c5f0cc71d9fbb2b605ddd.	2025-08-03 12:12:13 +08:00
yangdx	bfe6657b31	Remove auto_manage_storages_states option - Always manage storage states by LightRAG - Remove rag.initialize_storages() from all examples	2025-08-03 10:29:36 +08:00
yangdx	091f2b42c3	feat(performance): Optimize document deletion with entity/relation index - Introduces an index mapping documents to their corresponding entities and relations. This significantly speeds up `adelete_by_doc_id` by replacing slow graph traversal with a fast key-value lookup. - Refactors the ingestion pipeline (`merge_nodes_and_edges`) to populate this new index. Adds a one-time data migration script to backfill the index for existing data.	2025-08-03 09:19:02 +08:00
yangdx	32af45ff46	refactor: improve JSON parsing reliability with json-repair library Replace regex-based JSON extraction with json-repair for better handling of malformed LLM responses. Remove deprecated JSON parsing utilities and clean up keyword_extraction parameter across LLM providers. - Remove locate_json_string_body_from_string() and convert_response_to_json() - Use json-repair.loads() in extract_keywords_only() for robust parsing - Clean up LLM interfaces and remove unused parameters - Add json-repair dependency	2025-08-01 19:36:20 +08:00
yangdx	8271e1f6f1	Move OllamaServerInfos class to base module - Eliminate dependency of the core module on the API module.	2025-07-31 23:24:49 +08:00
yangdx	9d5603d35e	Set the default LLM temperature to 1.0 and centralize constant management	2025-07-31 17:15:10 +08:00
yangdx	c7bc4fc42c	Add track_id return to document processing pipeline	2025-07-30 10:27:12 +08:00
yangdx	cbaede8455	Add ScanResponse type for scan endpoint in webui	2025-07-30 03:11:09 +08:00
yangdx	7207598fc4	Fix track_id bugs and add track_id to scanning response	2025-07-30 03:06:20 +08:00
yangdx	93afa7d8a7	feat: add processing time tracking to document status with metadata field - Add metadata field to DocProcessingStatus with start_time and end_time tracking - Record processing timestamps using Unix time format (seconds precision) - Update all storage backends (JSON, MongoDB, Redis, PostgreSQL) for new field support - Maintain backward compatibility with default values for existing data - Add error_msg field for better error tracking during document processing	2025-07-29 23:42:33 +08:00
yangdx	6014b9bf73	feat: add track_id support for document processing progress monitoring - Add get_docs_by_track_id() method to all storage backends (MongoDB, PostgreSQL, Redis, JSON) - Implement automatic track_id generation with upload_/insert_ prefixes - Add /track_status/{track_id} API endpoint for frontend progress queries - Create database indexes for efficient track_id lookups - Enable real-time document processing status tracking across all storage types	2025-07-29 22:24:21 +08:00
yangdx	8274ed52d1	feat: separate document content from doc_status to improve performance This optimization significantly improves doc_status query/update performance by avoiding large string operations during frequent status checks.	2025-07-29 14:20:07 +08:00
yangdx	f2ffff063b	feat: refactor ollama server configuration management - Add ollama_server_infos attribute to LightRAG class with default initialization - Move default values to constants.py for centralized configuration - Refactor OllamaServerInfos class with property accessors and CLI support - Update OllamaAPI to get configuration through rag object instead of direct import - Add command line arguments for simulated model name and tag - Fix type imports to avoid circular dependencies	2025-07-28 01:38:35 +08:00
yangdx	598eecd06d	Refactor: Rename llm_model_max_token_size to summary_max_tokens This commit renames the parameter 'llm_model_max_token_size' to 'summary_max_tokens' for better clarity, as it specifically controls the token limit for entity relation summaries.	2025-07-28 00:49:08 +08:00
yangdx	ebaff228aa	feat: Add rerank score filtering with configurable threshold - Add DEFAULT_MIN_RERANK_SCORE constant (default: 0.0) - Add MIN_RERANK_SCORE environment variable support - Filter chunks with rerank scores below threshold in process_chunks_unified - Add info-level logging for filtering operations - Handle empty results gracefully after filtering - Maintain backward compatibility with non-reranked chunks	2025-07-27 16:37:44 +08:00
yangdx	b3c2987006	Reduce default MAX_TOKENS from 32000 to 10000	2025-07-26 08:13:49 +08:00
yangdx	983bacd87e	Update logger messages	2025-07-24 16:49:28 +08:00
yangdx	44b7ce222e	feat: add default storage dependencies and optimize imports - Add nano-vectordb and networkx to pyproject.toml dependencies - Replace dynamic imports with direct imports for 4 default storage implementations - Improve startup performance while maintaining backward compatibility	2025-07-24 16:14:26 +08:00
yangdx	05bc5cfb64	Improve task execution with early failure detection - Add early failure detection for async tasks - Cancel pending tasks on first exception	2025-07-19 10:14:22 +08:00
yangdx	5f7cb437e8	Centralize query parameters into LightRAG class This commit refactors query parameter management by consolidating settings like `top_k`, token limits, and thresholds into the `LightRAG` class, and consistently sourcing parameters from a single location.	2025-07-15 23:56:49 +08:00
yangdx	47341d3a71	Merge branch 'main' into rerank	2025-07-15 16:12:33 +08:00
yangdx	ccc2a20071	feat: remove deprecated MAX_TOKEN_SUMMARY parameter to prevent LLM output truncation - Remove MAX_TOKEN_SUMMARY parameter and related configurations - Eliminate forced token-based truncation in entity/relationship descriptions - Switch to fragment-count based summarization logic using FORCE_LLM_SUMMARY_ON_MERGE - Update FORCE_LLM_SUMMARY_ON_MERGE default from 6 to 4 for better summarization - Clean up documentation, environment examples, and API display code - Preserve backward compatibility by graceful parameter removal This change resolves issues where LLMs were forcibly truncating entity relationship descriptions mid-sentence, leading to incomplete and potentially inaccurate knowledge graph content. The new approach allows LLMs to generate complete descriptions while still providing summarization when multiple fragments need to be merged. Breaking Change: None - parameter removal is backward compatible Fixes: Entity relationship description truncation issues	2025-07-15 12:26:33 +08:00
zrguo	7c882313bb	remove chunk_rerank_top_k	2025-07-15 11:52:34 +08:00
yangdx	b03bb48e24	feat: Refine summary logic and add dedicated Ollama num_ctx config - Refactor the trigger condition for LLM-based summarization of entities and relations. Instead of relying on character length, the summary is now triggered when the number of merged description fragments exceeds a configured threshold. This provides a more robust and logical condition for consolidation. - Introduce the `OLLAMA_NUM_CTX` environment variable to explicitly configure the context window size (`num_ctx`) for Ollama models. This decouples the model's context length from the `MAX_TOKENS` parameter, which is now specifically used to limit input for summary generation, making the configuration clearer and more flexible. - Updated `README` files, `env.example`, and default values to reflect these changes.	2025-07-14 01:55:04 +08:00
yangdx	03b40937f7	Reduce embedding concurrency limit from 16 to 8	2025-07-13 03:13:52 +08:00
yangdx	39965d7ded	Move merging stage back controled by max parallel insert semhore	2025-07-12 03:32:08 +08:00
yangdx	3afdd1b67c	Fix initial count error for multi-process lock with key	2025-07-11 20:39:08 +08:00
yangdx	c47747da9e	Merge branch 'main' into merge_lock_with_key	2025-07-11 16:37:10 +08:00
yangdx	ef4870fda5	Combined entity and edge processing tasks and optimize merging with semaphore	2025-07-11 16:34:54 +08:00
yangdx	9aa2ed0837	Merge branch 'main' into rerank	2025-07-09 15:33:39 +08:00
yangdx	207f0a7f2a	Merge branch 'main' into merge_lock_with_key	2025-07-09 09:25:28 +08:00
yangdx	cb3bfc0e5b	Release semphore before merge stage	2025-07-09 09:24:44 +08:00
Anton Vice	b192f8c9a3	Fix: Handle NoneType error when processing documents without a file path The document processing pipeline would crash with a TypeError when a document was submitted as raw text via the API, as the file_path attribute would be None. This change adds a check to handle the None case gracefully, preventing the crash and allowing text-based documents to be indexed correctly.	2025-07-08 19:35:22 -03:00
zrguo	71cb3adb4f	Merge branch 'main' into rerank	2025-07-08 15:10:23 +08:00
yangdx	56d43de58a	Merge branch 'main' into merge_lock_with_key	2025-07-08 12:46:31 +08:00
zrguo	f5c80d7cde	Simplify Configuration	2025-07-08 11:16:34 +08:00
yangdx	9b7b2a9b0f	Reduce default embedding batch size from 32 to 10	2025-07-08 11:00:09 +08:00
zrguo	75dd4f3498	add rerank model	2025-07-07 22:44:59 +08:00
yangdx	ef79088f60	Move max_graph_nodes to global config	2025-07-07 21:53:57 +08:00
yangdx	033098c1bc	Feat: Add WORKSPACE support to all storage types	2025-07-07 00:57:21 +08:00
yangdx	1b2d295a4f	Remove namespace_prefix	2025-07-06 00:16:47 +08:00

1 2 3 4 5 ...

528 Commits