Machine Learning - Learning/Language Models

Machine Learning - Learning/Language Models ylai • Now • 100%

Scalable MatMul-free Language Modeling

arxiv.org

Machine Learning - Learning/Language Models ylai • Now • 100%

1-bit LLMs Could Solve AI’s Energy Demands

spectrum.ieee.org

Machine Learning - Learning/Language Models ylai • Now • 100%

Mistral 7B v0.2 Base (released at SHACK15sf hackathon)

github.com

GitHub: https://github.com/mistralai-sf24/hackathon \ X: https://twitter.com/MistralAILabs/status/1771670765521281370 >New release: Mistral 7B v0.2 Base (Raw pretrained model used to train Mistral-7B-Instruct-v0.2)\ >🔸 https://models.mistralcdn.com/mistral-7b-v0-2/mistral-7B-v0.2.tar \ >🔸 32k context window\ >🔸 Rope Theta = 1e6\ >🔸 No sliding window\ >🔸 How to fine-tune:

Machine Learning - Learning/Language Models ylai • Now • 50%

Evolving New Foundation Models: Unleashing the Power of Automating Model Development

sakana.ai

arXiv: https://arxiv.org/abs/2403.13187 \[cs.NE\]\ GitHub: https://github.com/SakanaAI/evolutionary-model-merge

Machine Learning - Learning/Language Models ylai • Now • 100%

GaLore: Advancing Large Model Training on Consumer-grade Hardware

huggingface.co

arXiv: https://arxiv.org/abs/2403.03507 [cs.LG]

Machine Learning - Learning/Language Models ylai • Now • 100%

How Chain-of-Thought Reasoning Helps Neural Networks Compute

www.quantamagazine.org

Machine Learning - Learning/Language Models ylai • Now • 100%

Why Are Large AI Models Being Red Teamed?

spectrum.ieee.org

Machine Learning - Learning/Language Models ylai • Now • 100%

GPT-4 won't run DOOM but will play the game poorly

www.theregister.com

Machine Learning - Learning/Language Models ylai • Now • 100%

LLMs become more covertly racist with human intervention

www.technologyreview.com

Machine Learning - Learning/Language Models ylai • Now • 100%

AI chatbot models ‘think’ in English even when using other languages

www.newscientist.com

Without paywall: https://archive.ph/Qq9Yd

Machine Learning - Learning/Language Models ylai • Now • 100%

AI Prompt Engineering Is Dead

spectrum.ieee.org

Machine Learning - Learning/Language Models ylai • Now • 100%

Large language models can do jaw-dropping things. But nobody knows exactly why.

www.technologyreview.com

Machine Learning - Learning/Language Models ylai • Now • 100%

Mixtral of Experts

https://arxiv.org/abs/2401.04088

Machine Learning - Learning/Language Models ylai • Now • 100%

Finetune LLMs on your own consumer hardware using tools from PyTorch and Hugging Face ecosystem

pytorch.org

Machine Learning - Learning/Language Models ylai • Now • 100%

“AI’s Ostensible Emergent Abilities Are a Mirage” paper won the Outstanding Paper Award at NeurIPS 2023

Previous Lemmy.ml post: https://lemmy.ml/post/1015476 Original X post (at Nitter): https://nitter.net/xwang_lk/status/1734356472606130646

Machine Learning - Learning/Language Models ylai • Now • 100%

High-Performance Llama 2 Training and Inference with PyTorch/XLA on Cloud TPUs

pytorch.org

Machine Learning - Learning/Language Models ylai • Now • 100%

Llemma: An Open Language Model For Mathematics

blog.eleuther.ai

Machine Learning - Learning/Language Models ylai • Now • 100%

“Large Language Models (in 2023)” (Talk by Hyung Won Chung, OpenAI, at Seoul National University)

www.youtube.com

Machine Learning - Learning/Language Models ylai • Now • 100%

LLM Finetuning Risks

https://llm-tuning-safety.github.io/

Machine Learning - Learning/Language Models ylai • Now • 100%

Phi 1.5 and the Shift Towards Smaller Models with Curated Data: A Closer Look

https://medium.com/ai-insights-cobet/phi-1-5-and-the-shift-towards-smaller-models-with-curated-data-a-closer-look-b7952a2e6730

Machine Learning - Learning/Language Models manitcor • Now • 100%

Stanford Online -Statistical Learning

https://www.youtube.com/playlist?list=PLoROMvodv4rOzrYsAxzQyHb8n_RWNuS1e

Machine Learning - Learning/Language Models manitcor • Now • 100%

Stanford University Lecture Collection | Convolutional Neural Networks

https://www.youtube.com/playlist?list=PL3FW7Lu3i5JvHM8ljYj-zLfQRF3EO8sYv

Machine Learning - Learning/Language Models manitcor • Now • 100%

Stanford CS229: Machine Learning

https://www.youtube.com/playlist?list=PLoROMvodv4rMiGQp3WXShtMGgzqpfVfbU

Machine Learning - Learning/Language Models manitcor • Now • 100%

MIT 6.S191: Introduction to Deep Learning

https://www.youtube.com/playlist?list=PLtBw6njQRU-rwp5__7C0oIVt26ZgjG9NI

Machine Learning - Learning/Language Models manitcor • Now • 100%

Carnegie Mellon University Deep Learning (11785 Fall 2022 Lectures)

https://www.youtube.com/playlist?list=PLp-0K3kfddPxRmjgjm0P1WT6H-gTqE8j9

Machine Learning - Learning/Language Models manitcor • Now • 100%

Applied Machine Learning (Cornell Tech CS 5787, Fall 2020)

https://www.youtube.com/playlist?list=PL2UML_KCiC0UlY7iCQDSiGDMovaupqc83

Machine Learning - Learning/Language Models manitcor • Now • 100%

DeepMind x UCL | Reinforcement Learning Course 2018

https://www.youtube.com/playlist?list=PLqYmG7hTraZBKeNJ-JE_eyJHZ7XgBoAyb

Machine Learning - Learning/Language Models ylai • Now • 100%

Chinchilla’s Death

https://espadrine.github.io/blog/posts/chinchilla-s-death.html

Machine Learning - Learning/Language Models ylai • Now • 100%

Is AI lying to us? These researchers built an LLM lie detector of sorts to find out

www.zdnet.com

Machine Learning - Learning/Language Models ylai • Now • 100%

Meet Mistral 7B, Mistral’s first LLM that beats Llama 2

dataconomy.com

Machine Learning - Learning/Language Models ylai • Now • 100%

Comparing Llama-2 and GPT-3 LLMs for HPC kernels generation

https://arxiv.org/abs/2309.07103

**Abstract** We evaluate the use of the open-source Llama-2 model for generating well-known, high-performance computing kernels (e.g., AXPY, GEMV, GEMM) on different parallel programming models and languages (e.g., C++: OpenMP, OpenMP Offload, OpenACC, CUDA, HIP; Fortran: OpenMP, OpenMP Offload, OpenACC; Python: numpy, Numba, pyCUDA, cuPy; and Julia: Threads, CUDA.jl, AMDGPU.jl). We built upon our previous work that is based on the OpenAI Codex, which is a descendant of GPT-3, to generate similar kernels with simple prompts via GitHub Copilot. Our goal is to compare the accuracy of Llama-2 and our original GPT-3 baseline by using a similar metric. Llama-2 has a simplified model that shows competitive or even superior accuracy. We also report on the differences between these foundational large language models as generative AI continues to redefine human-computer interactions. Overall, Copilot generates codes that are more reliable but less optimized, whereas codes generated by Llama-2 are less reliable but more optimized when correct.