LinkSoul/Chinese-LLaVA-Vision-Instructions
Viewer • Updated • 1.82M • 174 • 72
How to use amitha/mllava-baichuan2-zh with Transformers:
# Use a pipeline as a high-level helper
# Warning: Pipeline type "visual-question-answering" is no longer supported in transformers v5.
# You must load the model directly (see below) or downgrade to v4.x with:
# pip install "transformers<5.0.0"
from transformers import pipeline
pipe = pipeline("visual-question-answering", model="amitha/mllava-baichuan2-zh", trust_remote_code=True) # pip install -U transformers accelerate
# Load model directly
from transformers import AutoModelForVisualQuestionAnswering
model = AutoModelForVisualQuestionAnswering.from_pretrained("amitha/mllava-baichuan2-zh", trust_remote_code=True, device_map="auto")The Chinese Baichuan2-7B-Chat VLM trained via LORA for https://arxiv.org/abs/2406.11665.
The training data used for multimodal alignment and visual instruction tuning is from here.