Upload merged OLMoE-1B-7B with DoRA DPO

Browse files

Files changed (11) hide show

README.md +162 -0
chat_template.jinja +22 -0
config.json +32 -0
generation_config.json +6 -0
model-00001-of-00003.safetensors +3 -0
model-00002-of-00003.safetensors +3 -0
model-00003-of-00003.safetensors +3 -0
model.safetensors.index.json +0 -0
special_tokens_map.json +16 -0
tokenizer.json +0 -0
tokenizer_config.json +239 -0

README.md ADDED Viewed

	@@ -0,0 +1,162 @@

+---
+base_model: 1024m/OLMoE-1B-7B-0924-Base
+library_name: transformers
+pipeline_tag: text-generation
+license: apache-2.0
+tags:
+- dpo
+- dora
+- qlora
+- olmoe
+- alignment
+- preference-learning
+- merged
+datasets:
+- teknium/OpenHermes-2.5
+- HuggingFaceH4/ultrafeedback_binarized
+language:
+- en
+---
+# OLMoE-1B-7B DPO with DoRA (Merged)
+This is the **merged** version of [demonlxrd/olmoe-openhermes-ultrafeedback-dora-dpo](https://huggingface.co/demonlxrd/olmoe-openhermes-ultrafeedback-dora-dpo) - a preference-aligned OLMoE model trained with DoRA and DPO.
+## What's This?
+A fully merged model ready for production deployment. The DoRA adapter has been merged into the base [OLMoE-1B-7B](https://huggingface.co/1024m/OLMoE-1B-7B-0924-Base) weights for:
+- ✅ Faster inference (no adapter overhead)
+- ✅ vLLM compatibility
+- ✅ Simpler deployment
+- ✅ Production-ready
+Training pipeline:
+1. **SFT** on 20K examples from [OpenHermes-2.5](https://huggingface.co/datasets/teknium/OpenHermes-2.5)
+2. **DPO** on 10K preference pairs from [UltraFeedback](https://huggingface.co/datasets/HuggingFaceH4/ultrafeedback_binarized)
+3. **Merged** DoRA adapter into base weights
+## Quick Start
+### vLLM (Recommended)
+```bash
+# Serve
+vllm serve demonlxrd/olmoe-openhermes-ultrafeedback-dora-dpo-merged \
+  --max-model-len 4096 \
+  --dtype bfloat16
+# Inference
+curl -s http://localhost:8000/v1/chat/completions \
+  -H "Content-Type: application/json" \
+  -d '{
+    "model": "demonlxrd/olmoe-openhermes-ultrafeedback-dora-dpo-merged",
+    "messages": [
+      {"role": "user", "content": "Explain machine learning in simple terms."}
+    ],
+    "max_tokens": 200,
+    "temperature": 0.7
+  }' | jq -r '.choices[0].message.content'
+```
+### Python with Transformers
+```python
+import torch
+from transformers import AutoTokenizer, AutoModelForCausalLM
+tokenizer = AutoTokenizer.from_pretrained("demonlxrd/olmoe-openhermes-ultrafeedback-dora-dpo-merged")
+model = AutoModelForCausalLM.from_pretrained(
+    "demonlxrd/olmoe-openhermes-ultrafeedback-dora-dpo-merged",
+    device_map="auto",
+    torch_dtype=torch.bfloat16
+)
+messages = [{"role": "user", "content": "What is quantum computing?"}]
+prompt = tokenizer.apply_chat_template(messages, add_generation_prompt=True, tokenize=False)
+inputs = tokenizer(prompt, return_tensors="pt").to(model.device)
+with torch.inference_mode():
+    outputs = model.generate(**inputs, max_tokens=200, temperature=0.7)
+print(tokenizer.decode(outputs[0], skip_special_tokens=True))
+```
+### Python with OpenAI Client
+```python
+from openai import OpenAI
+client = OpenAI(
+    base_url="http://localhost:8000/v1",
+    api_key="dummy"
+)
+response = client.chat.completions.create(
+    model="demonlxrd/olmoe-openhermes-ultrafeedback-dora-dpo-merged",
+    messages=[
+        {"role": "user", "content": "Write a Python function to calculate fibonacci numbers."}
+    ],
+    max_tokens=300,
+    temperature=0.7
+)
+print(response.choices[0].message.content)
+```
+## Model Details
+| Parameter | Value |
+|-----------|-------|
+| Architecture | OLMoE (Mixture of Experts) |
+| Parameters | ~1B active, 7B total |
+| Precision | bfloat16 |
+| Context Length | 4096 tokens |
+| Training | SFT + DPO with DoRA adapters |
+| Base Model | 1024m/OLMoE-1B-7B-0924-Base |
+## Training Details
+- **Adapter Type**: DoRA (Weight-Decomposed LoRA)
+- **LoRA Rank**: 16
+- **Target Modules**: q_proj, v_proj
+- **Quantization during training**: 4-bit NF4
+- **DPO Beta**: 0.1
+- **Learning Rate**: 5e-5
+- **Hardware**: 2× NVIDIA A40 80GB
+## Chat Template
+```
+User:
+<message>
+Assistant:
+<response>
+```
+Roles supported: `system`, `user`, `assistant`
+## Why Use the Merged Version?
+- **Performance**: No adapter overhead during inference
+- **Compatibility**: Works with vLLM, TGI, and other optimized serving frameworks
+- **Simplicity**: Single model file, no need to load base + adapter separately
+- **Production-Ready**: Optimized for deployment at scale
+## Adapter Version
+Looking for the lightweight adapter weights? Check out [demonlxrd/olmoe-openhermes-ultrafeedback-dora-dpo](https://huggingface.co/demonlxrd/olmoe-openhermes-ultrafeedback-dora-dpo) (~8.3MB)
+## License
+Apache 2.0. Please also check the license of the [base model](https://huggingface.co/1024m/OLMoE-1B-7B-0924-Base).
+## Citation
+If you use this model, please cite:
+- **Base Model**: [1024m/OLMoE-1B-7B-0924-Base](https://huggingface.co/1024m/OLMoE-1B-7B-0924-Base)
+- **OpenHermes-2.5**: [teknium/OpenHermes-2.5](https://huggingface.co/datasets/teknium/OpenHermes-2.5)
+- **UltraFeedback**: [HuggingFaceH4/ultrafeedback_binarized](https://huggingface.co/datasets/HuggingFaceH4/ultrafeedback_binarized)
+- **TRL**: [HuggingFace TRL](https://github.com/huggingface/trl)

chat_template.jinja ADDED Viewed

	@@ -0,0 +1,22 @@

+    {% set sep = '\n\n' -%}
+    {% if bos_token is defined %}{{ bos_token }}{% endif -%}
+    {%- for m in messages -%}
+    {%- if m['role'] == 'system' -%}
+    System:
+    {{ m['content'] | trim }}{{ sep }}
+    {%- elif m['role'] == 'user' -%}
+    User:
+    {{ m['content'] | trim }}{{ sep }}
+    {%- elif m['role'] == 'assistant' -%}
+    Assistant:
+    {{ m['content'] | trim }}{{ sep }}
+    {%- elif m['role'] == 'tool' -%}
+    Tool:
+    {{ m['content'] | trim }}{{ sep }}
+    {%- endif -%}
+    {%- endfor -%}
+    {%- if add_generation_prompt -%}
+    Assistant:
+    {%- endif -%}

config.json ADDED Viewed

	@@ -0,0 +1,32 @@

+{
+  "architectures": [
+    "OlmoeForCausalLM"
+  ],
+  "attention_bias": false,
+  "attention_dropout": 0.0,
+  "clip_qkv": null,
+  "dtype": "bfloat16",
+  "eos_token_id": 50279,
+  "hidden_act": "silu",
+  "hidden_size": 2048,
+  "initializer_range": 0.02,
+  "intermediate_size": 1024,
+  "max_position_embeddings": 4096,
+  "model_type": "olmoe",
+  "norm_topk_prob": false,
+  "num_attention_heads": 16,
+  "num_experts": 64,
+  "num_experts_per_tok": 8,
+  "num_hidden_layers": 16,
+  "num_key_value_heads": 16,
+  "output_router_logits": false,
+  "pad_token_id": 1,
+  "rms_norm_eps": 1e-05,
+  "rope_scaling": null,
+  "rope_theta": 10000.0,
+  "router_aux_loss_coef": 0.01,
+  "tie_word_embeddings": false,
+  "transformers_version": "4.56.1",
+  "use_cache": true,
+  "vocab_size": 50304
+}

generation_config.json ADDED Viewed

	@@ -0,0 +1,6 @@

+{
+  "_from_model_config": true,
+  "eos_token_id": 50279,
+  "pad_token_id": 1,
+  "transformers_version": "4.56.1"
+}

model-00001-of-00003.safetensors ADDED Viewed

	@@ -0,0 +1,3 @@

+version https://git-lfs.github.com/spec/v1
+oid sha256:8f28aeaae685d658c68276a67e724ea261568b0bcd02fe08d663997c900a8f31
+size 4997744872

model-00002-of-00003.safetensors ADDED Viewed

	@@ -0,0 +1,3 @@

+version https://git-lfs.github.com/spec/v1
+oid sha256:49ec291a1b529d63c642776e8ca870614ccf9cff23de6efb1f32d1f5a403038e
+size 4997235176

model-00003-of-00003.safetensors ADDED Viewed

	@@ -0,0 +1,3 @@

+version https://git-lfs.github.com/spec/v1
+oid sha256:6f9c6c65aa1094abfd59b47854dfcdd8d00e6b25c6573189fb3d554ede876fa8
+size 3843741912

model.safetensors.index.json ADDED Viewed

The diff for this file is too large to render. See raw diff

special_tokens_map.json ADDED Viewed

	@@ -0,0 +1,16 @@

+{
+  "eos_token": {
+    "content": "<|endoftext|>",
+    "lstrip": false,
+    "normalized": false,
+    "rstrip": false,
+    "single_word": false
+  },
+  "pad_token": {
+    "content": "<|padding|>",
+    "lstrip": false,
+    "normalized": false,
+    "rstrip": false,
+    "single_word": false
+  }
+}

tokenizer.json ADDED Viewed

The diff for this file is too large to render. See raw diff

tokenizer_config.json ADDED Viewed

	@@ -0,0 +1,239 @@

+{
+  "add_bos_token": false,
+  "add_eos_token": false,
+  "add_prefix_space": false,
+  "added_tokens_decoder": {
+    "0": {
+      "content": "|||IP_ADDRESS|||",
+      "lstrip": false,
+      "normalized": true,
+      "rstrip": false,
+      "single_word": false,
+      "special": false
+    },
+    "1": {
+      "content": "<|padding|>",
+      "lstrip": false,
+      "normalized": false,
+      "rstrip": false,
+      "single_word": false,
+      "special": true
+    },
+    "50254": {
+      "content": "                        ",
+      "lstrip": false,
+      "normalized": true,
+      "rstrip": false,
+      "single_word": false,
+      "special": false
+    },
+    "50255": {
+      "content": "                       ",
+      "lstrip": false,
+      "normalized": true,
+      "rstrip": false,
+      "single_word": false,
+      "special": false
+    },
+    "50256": {
+      "content": "                      ",
+      "lstrip": false,
+      "normalized": true,
+      "rstrip": false,
+      "single_word": false,
+      "special": false
+    },
+    "50257": {
+      "content": "                     ",
+      "lstrip": false,
+      "normalized": true,
+      "rstrip": false,
+      "single_word": false,
+      "special": false
+    },
+    "50258": {
+      "content": "                    ",
+      "lstrip": false,
+      "normalized": true,
+      "rstrip": false,
+      "single_word": false,
+      "special": false
+    },
+    "50259": {
+      "content": "                   ",
+      "lstrip": false,
+      "normalized": true,
+      "rstrip": false,
+      "single_word": false,
+      "special": false
+    },
+    "50260": {
+      "content": "                  ",
+      "lstrip": false,
+      "normalized": true,
+      "rstrip": false,
+      "single_word": false,
+      "special": false
+    },
+    "50261": {
+      "content": "                 ",
+      "lstrip": false,
+      "normalized": true,
+      "rstrip": false,
+      "single_word": false,
+      "special": false
+    },
+    "50262": {
+      "content": "                ",
+      "lstrip": false,
+      "normalized": true,
+      "rstrip": false,
+      "single_word": false,
+      "special": false
+    },
+    "50263": {
+      "content": "               ",
+      "lstrip": false,
+      "normalized": true,
+      "rstrip": false,
+      "single_word": false,
+      "special": false
+    },
+    "50264": {
+      "content": "              ",
+      "lstrip": false,
+      "normalized": true,
+      "rstrip": false,
+      "single_word": false,
+      "special": false
+    },
+    "50265": {
+      "content": "             ",
+      "lstrip": false,
+      "normalized": true,
+      "rstrip": false,
+      "single_word": false,
+      "special": false
+    },
+    "50266": {
+      "content": "            ",
+      "lstrip": false,
+      "normalized": true,
+      "rstrip": false,
+      "single_word": false,
+      "special": false
+    },
+    "50267": {
+      "content": "           ",
+      "lstrip": false,
+      "normalized": true,
+      "rstrip": false,
+      "single_word": false,
+      "special": false
+    },
+    "50268": {
+      "content": "          ",
+      "lstrip": false,
+      "normalized": true,
+      "rstrip": false,
+      "single_word": false,
+      "special": false
+    },
+    "50269": {
+      "content": "         ",
+      "lstrip": false,
+      "normalized": true,
+      "rstrip": false,
+      "single_word": false,
+      "special": false
+    },
+    "50270": {
+      "content": "        ",
+      "lstrip": false,
+      "normalized": true,
+      "rstrip": false,
+      "single_word": false,
+      "special": false
+    },
+    "50271": {
+      "content": "       ",
+      "lstrip": false,
+      "normalized": true,
+      "rstrip": false,
+      "single_word": false,
+      "special": false
+    },
+    "50272": {
+      "content": "      ",
+      "lstrip": false,
+      "normalized": true,
+      "rstrip": false,
+      "single_word": false,
+      "special": false
+    },
+    "50273": {
+      "content": "     ",
+      "lstrip": false,
+      "normalized": true,
+      "rstrip": false,
+      "single_word": false,
+      "special": false
+    },
+    "50274": {
+      "content": "    ",
+      "lstrip": false,
+      "normalized": true,
+      "rstrip": false,
+      "single_word": false,
+      "special": false
+    },
+    "50275": {
+      "content": "   ",
+      "lstrip": false,
+      "normalized": true,
+      "rstrip": false,
+      "single_word": false,
+      "special": false
+    },
+    "50276": {
+      "content": "  ",
+      "lstrip": false,
+      "normalized": true,
+      "rstrip": false,
+      "single_word": false,
+      "special": false
+    },
+    "50277": {
+      "content": "|||EMAIL_ADDRESS|||",
+      "lstrip": false,
+      "normalized": true,
+      "rstrip": false,
+      "single_word": false,
+      "special": false
+    },
+    "50278": {
+      "content": "|||PHONE_NUMBER|||",
+      "lstrip": false,
+      "normalized": true,
+      "rstrip": false,
+      "single_word": false,
+      "special": false
+    },
+    "50279": {
+      "content": "<|endoftext|>",
+      "lstrip": false,
+      "normalized": false,
+      "rstrip": false,
+      "single_word": false,
+      "special": true
+    }
+  },
+  "bos_token": null,
+  "clean_up_tokenization_spaces": true,
+  "eos_token": "<|endoftext|>",
+  "extra_special_tokens": {},
+  "model_max_length": 1000000000000000019884624838656,
+  "pad_token": "<|padding|>",
+  "tokenizer_class": "GPTNeoXTokenizer",
+  "unk_token": null
+}