--- base_model: - Qwen/Qwen2.5-7B-Instruct language: - en - zh license: mit pipeline_tag: question-answering library_name: transformers tags: - biology - finance - text-generation-inference --- # HierSearch: A Hierarchical Enterprise Deep Search Framework Integrating Local and Web Searches ## Model Information We release the agent model used in **HierSearch: A Hierarchical Enterprise Deep Search Framework Integrating Local and Web Searches**.

Useful links: 📝 Paper (arXiv) • 🤗 Paper (Hugging Face) • 🧩 Github

1. We explore the deep search framework in multi-knowledge-source scenarios and propose a hierarchical agentic paradigm and train with HRL; 2. We notice drawbacks of the naive information transmission among deep search agents and developed a knowledge refiner suitable for multi-knowledge-source scenarios; 3. Our proposed approach for reliable and effective deep search across multiple knowledge sources outperforms existing baselines the flat-RL solution in various domains. 🌹 If you use this model, please ✨star our **[GitHub repository](https://github.com/plageon/HierSearch)** or upvote our **[paper](https://huggingface.co/papers/2508.08088)** to support us. Your star means a lot! ## Sample Usage You can load and use this model directly with the Hugging Face `transformers` library for basic text generation or question-answering inference. For the full HierSearch framework capabilities, please refer to the [official GitHub repository](https://github.com/plageon/HierSearch). ```python from transformers import AutoModelForCausalLM, AutoTokenizer import torch model_id = "zstanjj/HierSearch-Planner-Agent" # This model represents the Planner Agent. # Other agent models include "zstanjj/HierSearch-Local-Agent" or "zstanjj/HierSearch-Web-Agent". tokenizer = AutoTokenizer.from_pretrained(model_id) model = AutoModelForCausalLM.from_pretrained( model_id, torch_dtype=torch.bfloat16, # Or torch.float16 depending on your hardware device_map="auto" # Or specify your device, e.g., "cuda:0" ) # Example for a question-answering interaction with the Planner Agent messages = [ {"role": "user", "content": "Explain the concept of Hierarchical Reinforcement Learning as applied in this paper."}, ] # Apply chat template and tokenize inputs text = tokenizer.apply_chat_template( messages, tokenize=False, add_generation_prompt=True ) model_inputs = tokenizer([text], return_tensors="pt").to(model.device) # Generate response generated_ids = model.generate( model_inputs.input_ids, max_new_tokens=1024, # Adjust max_new_tokens as needed for detailed answers temperature=0.7, # Adjust generation parameters for diversity do_sample=True, eos_token_id=tokenizer.eos_token_id, # Ensure generation stops at EOS token pad_token_id=tokenizer.pad_token_id # Set pad_token_id for proper generation ) # Decode and print the output decoded_output = tokenizer.batch_decode(generated_ids, skip_special_tokens=True)[0] print(decoded_output) ```