Hugging Face
Models
Datasets
Spaces
Buckets
new
Docs
Enterprise
Pricing
Website
Tasks
HuggingChat
Collections
Languages
Organizations
Community
Blog
Posts
Daily Papers
Hardware
Learn
Discord
Forum
GitHub
Solutions
Team & Enterprise
Hugging Face PRO
Enterprise Support
Inference Providers
Inference Endpoints
Storage Buckets
Log In
Sign Up
🤝
Open to Collab
358.3
TFLOPS
AbstractPhila
PRO
AbstractPhil
19
6
27
Follow
liujiahiu's profile picture
Pq234's profile picture
RandyXia's profile picture
94 followers
·
128 following
https://civitai.com/user/AbstractPhila
AbstractEyes
AI & ML interests
datasets, research papers, experimentation, vision, classification, text encoders, tokenization, llms, diffusion, distillation, and more.
Recent Activity
posted
an
update
about 5 hours ago
Say hello to the AlephLLM: Mini-Beatrix - in her huggingface space! She is currently stepped at 8000 steps in the first couple datasets, so she's not very smart yet. https://huggingface.co/spaces/AbstractPhil/alephllm-chat Be warned, whatever you say will be recorded in a public cache. The AlephLLM prototype is currently in full training with SDPA attention. https://huggingface.co/AbstractPhil/alephllm-mini-beatrix-training https://github.com/AbstractEyes/alephllm Here's the model code and training code for the prototype. As the training progresses, the AlephLLM will become more coherent and communicative, the tensorboard will consist of a large series of useful and useless analysis, and each checkpoint recorded at around 2000 steps unless the train crashes or the system faults. It will take about 9 hours for the first few datasets to converge, then I'll train a chat AMOE expert cluster to see if she wants to speak yet. Until then, she's learning. Yes I know it's early, but there isn't much more I could think of to analyze the AlephLM directly currently. The only way train the AlephLLM, is to train the full AlephLLM prototype. The bigger training has to run, otherwise the analysis won't matter. As it progresses, the analysis and huge amount of tensorboard statistics will flood out. Everything is transparent through the process from start to finish, everything recorded.
updated
a model
about 5 hours ago
AbstractPhil/alephllm-mini-beatrix-training
replied
to
their
post
about 7 hours ago
The upcoming AlephLM LLM prototype "Mini-Beatrix" is based on protocols, rules, and laws established through the process of training AlephLM systems. This will be a first attempt at a smaller full pretrain/finetune of the AlephLLM on raw data, and this will require over a billion unigram tokens. Mini-Beatrix will inherit an appropriately adapted AlephLM MOE structure containing a multitude of trained experts, a gating system, a long context RoPE system, MHA attention, and a series of hypothesis to answer upon Mini-Beatrix's pretrain and finetune completion. While focusing on resolving corruptions and invalidity possibly present in the splat attention, the solutions raised SDPA attention protocol token recall ceiling from 0.91 to 0.993. With that the splat attention raised from 0.81 to 0.89~ splat being around 3x the speed is still imperfect. So far so good. The corruptions have resolved multiple core component overlapping problems causing the AlephLM's inability to handle the trigram system, the structure of the SVAE having faulty trigram structures, and additionally a multitude of other systems in the lineup that were inheriting the corruptions from the core experiment sets. These corruptions resolved show that the accuracy of standard multiheaded attention will provide the necessary token recall for full LM capacity, and with that If and WHEN I solve the Rorschach Splat attention will be the faster alternative at >=r1 0.99%, only then. The splat attention's considerably larger head count still contains unresolved inconsistencies. That being said the SDPA MHA attention will be present for the first attempted mini-llm train, which will be named "Mini-Beatrix" with the appropriate sizing associated with this. The only thing that will change Mini-Beatrix's trajectory will be if Splat attention is perfected between today and next week, which will likely take longer unless I run into a core corruption that has been overlooked through hundreds of analysis.
View all activity
Organizations
AbstractPhil
's datasets
83
Sort: Recently updated
AbstractPhil/alephllm-chat-history
Viewer
•
Updated
about 7 hours ago
•
1
AbstractPhil/captionbert-8192-v2-consensus
Updated
11 days ago
•
202
AbstractPhil/conceptual-captions-12m-webdataset-berts
Viewer
•
Updated
11 days ago
•
32.3M
•
534
•
1
AbstractPhil/bulk-cc12m-features
Viewer
•
Updated
13 days ago
•
121M
•
3.86k
AbstractPhil/tower-probes-results
Viewer
•
Updated
22 days ago
•
17
•
142
AbstractPhil/qwen-deepfashion-fused
Viewer
•
Updated
Jul 12
•
122k
•
1.53k
•
1
AbstractPhil/qwen-synth-characters-fused
Viewer
•
Updated
Jul 10
•
42.7k
•
1.73k
AbstractPhil/qwen-synth-characters-100-json-test
Viewer
•
Updated
Jul 10
•
1k
•
54
AbstractPhil/anima-brent-90k-cache
Updated
Jul 5
•
48
AbstractPhil/qwen-synth-characters
Viewer
•
Updated
Jul 3
•
61k
•
122
AbstractPhil/qwen-deepfashion
Viewer
•
Updated
Jul 3
•
160k
•
202
AbstractPhil/diffusion-pipe-cache-test1
Viewer
•
Updated
Jun 27
•
8.92k
•
68
AbstractPhil/anima-90k-cache
Updated
Jun 26
•
84
AbstractPhil/diffusion-pretrain-set-ft1
Viewer
•
Updated
Jun 23
•
1.46M
•
1.63k
•
1
AbstractPhil/diffusion-pretrain-set-ft1-1024
Viewer
•
Updated
Jun 11
•
1.14M
•
716
AbstractPhil/sdxl-qwen-phase1-cache
Viewer
•
Updated
Jun 6
•
86k
•
323
AbstractPhil/geolip-sdxl-fid-scoring
Viewer
•
Updated
Jun 5
•
2.8k
•
101
AbstractPhil/sdxl-qwen-phase0
Viewer
•
Updated
Jun 4
•
86k
•
266
•
3
AbstractPhil/IMDB-PUBLIC-SCRAPED
Preview
•
Updated
May 19
•
108
•
1
AbstractPhil/ldhnam-deepfashion_controlnet
Viewer
•
Updated
May 19
•
26k
•
25
AbstractPhil/ffhq_flux_latents_repaired
Viewer
•
Updated
May 19
•
40.8k
•
223
AbstractPhil/synthetic-characters
Viewer
•
Updated
May 19
•
149k
•
349
AbstractPhil/CN_pose3D_V10_512
Viewer
•
Updated
May 19
•
66.5k
•
70
AbstractPhil/CN_pose3D_V7_512
Viewer
•
Updated
May 19
•
255k
•
479
AbstractPhil/synthetic-object-relations-json
Viewer
•
Updated
May 18
•
5k
•
43
AbstractPhil/cc-task1-json
Preview
•
Updated
May 18
•
46
AbstractPhil/cc-prompts-sharded
Viewer
•
Updated
May 15
•
3.32M
•
10
AbstractPhil/json-coco-format
Viewer
•
Updated
May 14
•
129k
•
195
AbstractPhil/svae-freckles-4096-cifar10
Viewer
•
Updated
Apr 10
•
60k
•
61
AbstractPhil/ryan-spearman-prepared-features
Viewer
•
Updated
Mar 27
•
1
•
68
Previous
1
2
3
Next