Text Generation
Transformers
PyTorch
Safetensors
English
llama
text-generation-inference
unsloth
trl
sft
conversational
Instructions to use kparkhade/Llama-3.1-8B with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use kparkhade/Llama-3.1-8B with Transformers:
# Use a pipeline as a high-level helper from transformers import pipeline pipe = pipeline("text-generation", model="kparkhade/Llama-3.1-8B") messages = [ {"role": "user", "content": "Who are you?"}, ] pipe(messages)# Load model directly from transformers import AutoTokenizer, AutoModelForCausalLM tokenizer = AutoTokenizer.from_pretrained("kparkhade/Llama-3.1-8B") model = AutoModelForCausalLM.from_pretrained("kparkhade/Llama-3.1-8B", device_map="auto") messages = [ {"role": "user", "content": "Who are you?"}, ] inputs = tokenizer.apply_chat_template( messages, add_generation_prompt=True, tokenize=True, return_dict=True, return_tensors="pt", ).to(model.device) outputs = model.generate(**inputs, max_new_tokens=40) print(tokenizer.decode(outputs[0][inputs["input_ids"].shape[-1]:])) - Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- vLLM
How to use kparkhade/Llama-3.1-8B with vLLM:
Install from pip and serve model
# Install vLLM from pip: pip install vllm # Start the vLLM server: vllm serve "kparkhade/Llama-3.1-8B" # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:8000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "kparkhade/Llama-3.1-8B", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }'Use Docker
docker model run hf.co/kparkhade/Llama-3.1-8B
- SGLang
How to use kparkhade/Llama-3.1-8B with SGLang:
Install from pip and serve model
# Install SGLang from pip: pip install sglang # Start the SGLang server: python3 -m sglang.launch_server \ --model-path "kparkhade/Llama-3.1-8B" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "kparkhade/Llama-3.1-8B", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }'Use Docker images
docker run --gpus all \ --shm-size 32g \ -p 30000:30000 \ -v ~/.cache/huggingface:/root/.cache/huggingface \ --env "HF_TOKEN=<secret>" \ --ipc=host \ lmsysorg/sglang:latest \ python3 -m sglang.launch_server \ --model-path "kparkhade/Llama-3.1-8B" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "kparkhade/Llama-3.1-8B", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }' - Unsloth Studio
How to use kparkhade/Llama-3.1-8B with Unsloth Studio:
Install Unsloth Studio (macOS, Linux, WSL)
curl -fsSL https://unsloth.ai/install.sh | sh # Run unsloth studio unsloth studio -H 0.0.0.0 -p 8888 # Then open http://localhost:8888 in your browser # Search for kparkhade/Llama-3.1-8B to start chatting
Install Unsloth Studio (Windows)
irm https://unsloth.ai/install.ps1 | iex # Run unsloth studio unsloth studio -H 0.0.0.0 -p 8888 # Then open http://localhost:8888 in your browser # Search for kparkhade/Llama-3.1-8B to start chatting
Using HuggingFace Spaces for Unsloth
# No setup required # Open https://huggingface.co/spaces/unsloth/studio in your browser # Search for kparkhade/Llama-3.1-8B to start chatting
Load model with FastModel
pip install unsloth from unsloth import FastModel model, tokenizer = FastModel.from_pretrained( model_name="kparkhade/Llama-3.1-8B", max_seq_length=2048, ) - Docker Model Runner
How to use kparkhade/Llama-3.1-8B with Docker Model Runner:
docker model run hf.co/kparkhade/Llama-3.1-8B
Uploaded model: Llama 3.1 8B Finetuned
- Developed by: kparkhade
- License: apache-2.0
- Base model : unsloth/Meta-Llama-3.1-8B-bnb-4bit
Overview
This fine-tuned Llama 3.1 8B model was optimized for efficient text generation tasks. By leveraging advanced optimization techniques from Unsloth and Hugging Face's TRL library, training was completed 2x faster than conventional methods.
Key Features
- Speed Optimized: Training was accelerated with the Unsloth framework, significantly reducing resource consumption.
- Model Compatibility: Compatible with Hugging Face's ecosystem for seamless integration.
- Quantization: Built on a 4-bit quantized base model for efficient deployment and inference.
Usage
from transformers import AutoModelForCausalLM, AutoTokenizer
# Load the model and tokenizer
model_name = "kparkhade/Llama-3.1-8B"
model = AutoModelForCausalLM.from_pretrained(model_name)
tokenizer = AutoTokenizer.from_pretrained(model_name)
# Generate text
inputs = tokenizer("Your input prompt here", return_tensors="pt")
outputs = model.generate(**inputs)
print(tokenizer.decode(outputs[0], skip_special_tokens=True))
Applications
This model can be used for:
- Creative writing (e.g., story or poetry generation)
- Generating conversational responses
- Assisting with coding-related queries
Acknowledgements
Special thanks to the Unsloth team for providing tools that make model fine-tuning faster and more efficient.
- Downloads last month
- 3
Model tree for kparkhade/Llama-3.1-8B
Base model
meta-llama/Llama-3.1-8B Quantized
unsloth/Meta-Llama-3.1-8B-bnb-4bit