<?xml version="1.0" encoding="UTF-8"?>
<feed xmlns="http://www.w3.org/2005/Atom" xml:lang="en">
    <title>Brenner Cruvinel - ai safety</title>
    <subtitle>AI researcher and product designer building at the edge of computer science and mental health.</subtitle>
    <link rel="self" type="application/atom+xml" href="https://brennercruvinel.blog/tags/ai-safety/atom.xml"/>
    <link rel="alternate" type="text/html" href="https://brennercruvinel.blog"/>
    <generator uri="https://www.getzola.org/">Zola</generator>
    <updated>2026-04-08T00:00:00+00:00</updated>
    <id>https://brennercruvinel.blog/tags/ai-safety/atom.xml</id>
    <entry xml:lang="en">
        <title>Model Extraction Attacks contra Redes Neurais Black-Box (2020-2026)</title>
        <published>2026-04-08T00:00:00+00:00</published>
        <updated>2026-04-08T00:00:00+00:00</updated>
        
        <author>
          <name>
            Brenner Cruvinel
          </name>
        </author>
        
        <link rel="alternate" type="text/html" href="https://brennercruvinel.blog/blog/model-extraction-attacks/"/>
        <id>https://brennercruvinel.blog/blog/model-extraction-attacks/</id>
        
        <content type="html" xml:base="https://brennercruvinel.blog/blog/model-extraction-attacks/">&lt;h3 id=&quot;model-extraction-attacks-on-black-box-neural-networks-2020-2026&quot;&gt;Model Extraction Attacks on Black-Box Neural Networks (2020-2026)&lt;&#x2F;h3&gt;
&lt;p&gt;monorep fresh co, ataques de extracao contra modelos black-box, surveys e literatura agregada.&lt;&#x2F;p&gt;
&lt;ol&gt;
&lt;li&gt;B3: Backdoor Attacks against Black-box Machine Learning Models | ACM Transactions on Privacy and Security - https:&#x2F;&#x2F;dl.acm.org&#x2F;doi&#x2F;10.1145&#x2F;3605212&lt;&#x2F;li&gt;
&lt;li&gt;[PDF] CloudLeak: Large-Scale Deep Learning Models Stealing Through Adversarial Examples | Semantic Scholar - https:&#x2F;&#x2F;www.semanticscholar.org&#x2F;paper&#x2F;CloudLeak:-Large-Scale-Deep-Learning-Models-Through-Yu-Yang&#x2F;4d548fd21aad60e3052455e22b7a57cc1f06e3c3&lt;&#x2F;li&gt;
&lt;li&gt;Stealing Neural Networks With Model Extraction Attacks, Dr. Nicholas Carlini | Commonwealth Cyber Initiative (CCI) | Virginia Tech - https:&#x2F;&#x2F;cyberinitiative.org&#x2F;events-programs&#x2F;2020&#x2F;stealing-neural-networks-with-model-extraction-attacks-dr-nich.html&lt;&#x2F;li&gt;
&lt;li&gt;I Know What You Trained Last Summer: A Survey on Stealing Machine Learning Models and Defences | ACM Computing Surveys - https:&#x2F;&#x2F;dl.acm.org&#x2F;doi&#x2F;10.1145&#x2F;3595292&lt;&#x2F;li&gt;
&lt;li&gt;GitHub - kzhao5&#x2F;Model-Extraction-Stealing-Attacks-Machine-Learning-Literature - https:&#x2F;&#x2F;github.com&#x2F;kzhao5&#x2F;Model-Extraction-Stealing-Attacks-Machine-Learning-Literature&lt;&#x2F;li&gt;
&lt;li&gt;First Model-Stealing Attack Reveals Secrets of Black-Box Production Language Models | Synced - https:&#x2F;&#x2F;syncedreview.com&#x2F;2024&#x2F;03&#x2F;27&#x2F;first-model-stealing-attack-reveals-secrets-of-black-box-production-language-models&#x2F;&lt;&#x2F;li&gt;
&lt;li&gt;A realistic model extraction attack against graph neural networks - ScienceDirect - https:&#x2F;&#x2F;www.sciencedirect.com&#x2F;science&#x2F;article&#x2F;abs&#x2F;pii&#x2F;S0950705124007780&lt;&#x2F;li&gt;
&lt;li&gt;[1602.02697] Practical Black-Box Attacks against Machine Learning - https:&#x2F;&#x2F;arxiv.org&#x2F;abs&#x2F;1602.02697&lt;&#x2F;li&gt;
&lt;li&gt;Practical Black-Box Attacks against Machine Learning | Proceedings of the 2017 ACM on Asia Conference on Computer and Communications Security - https:&#x2F;&#x2F;dl.acm.org&#x2F;doi&#x2F;10.1145&#x2F;3052973.3053009&lt;&#x2F;li&gt;
&lt;li&gt;Black-box targeted adversarial attacks for deep neural networks - ScienceDirect - https:&#x2F;&#x2F;www.sciencedirect.com&#x2F;science&#x2F;article&#x2F;abs&#x2F;pii&#x2F;S0925231225016893&lt;&#x2F;li&gt;
&lt;&#x2F;ol&gt;
&lt;h3 id=&quot;stealing-machine-learning-models-via-prediction-apis-arxiv&quot;&gt;Stealing Machine Learning Models via Prediction APIs (arxiv)&lt;&#x2F;h3&gt;
&lt;p&gt;o paper seminal do Tramer e suas variantes, mais codigo de extracao via prediction API.&lt;&#x2F;p&gt;
&lt;ol&gt;
&lt;li&gt;[1609.02943] Stealing Machine Learning Models via Prediction APIs - https:&#x2F;&#x2F;arxiv.org&#x2F;abs&#x2F;1609.02943&lt;&#x2F;li&gt;
&lt;li&gt;Stealing Machine Learning Models via Prediction APIs - https:&#x2F;&#x2F;www.usenix.org&#x2F;system&#x2F;files&#x2F;conference&#x2F;usenixsecurity16&#x2F;sec16_paper_tramer.pdf&lt;&#x2F;li&gt;
&lt;li&gt;Stealing machine learning models via prediction APIs | Proceedings of the 25th USENIX Conference on Security Symposium - https:&#x2F;&#x2F;dl.acm.org&#x2F;doi&#x2F;10.5555&#x2F;3241094.3241142&lt;&#x2F;li&gt;
&lt;li&gt;[1609.02943] Stealing Machine Learning Models via Prediction APIs (HTML version) - https:&#x2F;&#x2F;ar5iv.labs.arxiv.org&#x2F;html&#x2F;1609.02943&lt;&#x2F;li&gt;
&lt;li&gt;[1609.02943v2] Stealing Machine Learning Models via Prediction APIs (Updated version) - https:&#x2F;&#x2F;arxiv.org&#x2F;abs&#x2F;1609.02943v2&lt;&#x2F;li&gt;
&lt;li&gt;Efficient and Effective Model Extraction - https:&#x2F;&#x2F;arxiv.org&#x2F;html&#x2F;2409.14122v1&lt;&#x2F;li&gt;
&lt;li&gt;Scholars@Duke publication: Stealing Machine Learning Models via Prediction APIs - https:&#x2F;&#x2F;scholars.duke.edu&#x2F;publication&#x2F;1493919&lt;&#x2F;li&gt;
&lt;li&gt;GitHub - ftramer&#x2F;Steal-ML: Model extraction attacks on Machine-Learning-as-a-Service platforms - https:&#x2F;&#x2F;github.com&#x2F;ftramer&#x2F;Steal-ML&lt;&#x2F;li&gt;
&lt;li&gt;[PDF] Stealing Machine Learning Models via Prediction APIs | Semantic Scholar - https:&#x2F;&#x2F;www.semanticscholar.org&#x2F;paper&#x2F;Stealing-Machine-Learning-Models-via-Prediction-Tram%C3%A8r-Zhang&#x2F;d15b21bdd117877c2d0e865b17d6a336737aea99&lt;&#x2F;li&gt;
&lt;li&gt;Stealing Machine Learning Models via Prediction APIs | USENIX - https:&#x2F;&#x2F;www.usenix.org&#x2F;conference&#x2F;usenixsecurity16&#x2F;technical-sessions&#x2F;presentation&#x2F;tramer&lt;&#x2F;li&gt;
&lt;&#x2F;ol&gt;
&lt;h3 id=&quot;model-inversion-attacks-2020-2025-deep-leakage-from-gradients&quot;&gt;model inversion attacks 2020-2025 (deep leakage from gradients)&lt;&#x2F;h3&gt;
&lt;p&gt;inversao de modelo e vazamento por gradiente, com foco em federated learning e training de language models.&lt;&#x2F;p&gt;
&lt;ol&gt;
&lt;li&gt;Iterative and mixed-spaces image gradient inversion attack in federated learning | Cybersecurity - https:&#x2F;&#x2F;cybersecurity.springeropen.com&#x2F;articles&#x2F;10.1186&#x2F;s42400-024-00227-7&lt;&#x2F;li&gt;
&lt;li&gt;Understanding Deep Gradient Leakage via Inversion Influence Functions - PubMed - https:&#x2F;&#x2F;pubmed.ncbi.nlm.nih.gov&#x2F;38606303&#x2F;&lt;&#x2F;li&gt;
&lt;li&gt;Deep leakage from gradients | Proceedings of the 33rd International Conference on Neural Information Processing Systems - https:&#x2F;&#x2F;dl.acm.org&#x2F;doi&#x2F;10.5555&#x2F;3454287.3455610&lt;&#x2F;li&gt;
&lt;li&gt;Wasserstein Distance-Based Deep Leakage from Gradients - PMC - https:&#x2F;&#x2F;www.ncbi.nlm.nih.gov&#x2F;pmc&#x2F;articles&#x2F;PMC10217429&#x2F;&lt;&#x2F;li&gt;
&lt;li&gt;Uncovering Gradient Inversion Risks in Practical Language Model Training | Proceedings of the 2024 on ACM SIGSAC Conference on Computer and Communications Security - https:&#x2F;&#x2F;dl.acm.org&#x2F;doi&#x2F;10.1145&#x2F;3658644.3690292&lt;&#x2F;li&gt;
&lt;li&gt;Deep Leakage from Gradients | SpringerLink - https:&#x2F;&#x2F;link.springer.com&#x2F;chapter&#x2F;10.1007&#x2F;978-3-030-63076-8_2&lt;&#x2F;li&gt;
&lt;li&gt;Understanding Deep Gradient Leakage via Inversion Influence Functions - https:&#x2F;&#x2F;arxiv.org&#x2F;html&#x2F;2309.13016&lt;&#x2F;li&gt;
&lt;li&gt;[1906.08935] Deep Leakage from Gradients - https:&#x2F;&#x2F;arxiv.org&#x2F;abs&#x2F;1906.08935&lt;&#x2F;li&gt;
&lt;li&gt;Inverting Gradient Attacks Naturally Makes Data Poisons: An Availability Attack on Neural Networks - https:&#x2F;&#x2F;arxiv.org&#x2F;html&#x2F;2410.21453v1&lt;&#x2F;li&gt;
&lt;li&gt;A Survey of the Implementations of Model Inversion Attacks | SpringerLink - https:&#x2F;&#x2F;link.springer.com&#x2F;chapter&#x2F;10.1007&#x2F;978-3-031-30648-8_1&lt;&#x2F;li&gt;
&lt;&#x2F;ol&gt;
&lt;h3 id=&quot;onnx-safetensors-serializacao-de-pesos&quot;&gt;ONNX, SafeTensors, serializacao de pesos&lt;&#x2F;h3&gt;
&lt;p&gt;formatos de armazenamento de peso de rede neural e conversao entre eles. relevante pra entender superficie de ataque na serializacao.&lt;&#x2F;p&gt;
&lt;ol&gt;
&lt;li&gt;ONNX - https:&#x2F;&#x2F;huggingface.co&#x2F;docs&#x2F;transformers&#x2F;en&#x2F;serialization&lt;&#x2F;li&gt;
&lt;li&gt;Safetensors, CKPT, ONNX, GGUF, and Other Key AI Model Formats | NEDNEX - https:&#x2F;&#x2F;nednex.com&#x2F;en&#x2F;what-are-safetensors&#x2F;&lt;&#x2F;li&gt;
&lt;li&gt;Model Weights File Formats in Machine Learning - https:&#x2F;&#x2F;learnopencv.com&#x2F;model-weights-file-formats-in-machine-learning&#x2F;&lt;&#x2F;li&gt;
&lt;li&gt;Export to ONNX - https:&#x2F;&#x2F;huggingface.co&#x2F;docs&#x2F;transformers&#x2F;v4.36.1&#x2F;serialization&lt;&#x2F;li&gt;
&lt;li&gt;justinchuby&#x2F;onnx-safetensors | DeepWiki - https:&#x2F;&#x2F;deepwiki.com&#x2F;justinchuby&#x2F;onnx-safetensors&lt;&#x2F;li&gt;
&lt;li&gt;Common AI Model Formats - https:&#x2F;&#x2F;huggingface.co&#x2F;blog&#x2F;ngxson&#x2F;common-ai-model-formats&lt;&#x2F;li&gt;
&lt;li&gt;python - How to convert safetensors model to onnx model? - Stack Overflow - https:&#x2F;&#x2F;stackoverflow.com&#x2F;questions&#x2F;77855742&#x2F;how-to-convert-safetensors-model-to-onnx-model&lt;&#x2F;li&gt;
&lt;li&gt;GitHub - justinchuby&#x2F;onnx-safetensors: Use safetensors with ONNX - https:&#x2F;&#x2F;github.com&#x2F;justinchuby&#x2F;onnx-safetensors&lt;&#x2F;li&gt;
&lt;li&gt;Save external data as safetensors - onnx&#x2F;onnx - Discussion #5461 - https:&#x2F;&#x2F;github.com&#x2F;onnx&#x2F;onnx&#x2F;discussions&#x2F;5461&lt;&#x2F;li&gt;
&lt;li&gt;GitHub - huggingface&#x2F;safetensors: Simple, safe way to store and distribute tensors - https:&#x2F;&#x2F;github.com&#x2F;huggingface&#x2F;safetensors&lt;&#x2F;li&gt;
&lt;&#x2F;ol&gt;
&lt;h3 id=&quot;distributed-training-e-sharding-fsdp&quot;&gt;distributed training e sharding (FSDP)&lt;&#x2F;h3&gt;
&lt;p&gt;tecnicas de sharding de modelo em treino distribuido, FSDP no centro.&lt;&#x2F;p&gt;
&lt;ol&gt;
&lt;li&gt;Fully Sharded Data Parallel: faster AI training with fewer GPUs Engineering at Meta - https:&#x2F;&#x2F;engineering.fb.com&#x2F;2021&#x2F;07&#x2F;15&#x2F;open-source&#x2F;fsdp&#x2F;&lt;&#x2F;li&gt;
&lt;li&gt;
&lt;ol start=&quot;2&quot;&gt;
&lt;li&gt;Sharding, Model Parallelism on the IPU with TensorFlow: Sharding and Pipelining - https:&#x2F;&#x2F;docs.graphcore.ai&#x2F;projects&#x2F;tf-model-parallelism&#x2F;en&#x2F;latest&#x2F;sharding.html&lt;&#x2F;li&gt;
&lt;&#x2F;ol&gt;
&lt;&#x2F;li&gt;
&lt;li&gt;Sharded Data Parallelism - Amazon SageMaker AI - https:&#x2F;&#x2F;docs.aws.amazon.com&#x2F;sagemaker&#x2F;latest&#x2F;dg&#x2F;model-parallel-extended-features-pytorch-sharded-data-parallelism.html&lt;&#x2F;li&gt;
&lt;li&gt;[2305.01868] Pre-train and Search: Efficient Embedding Table Sharding with Pre-trained Neural Cost Models - https:&#x2F;&#x2F;arxiv.org&#x2F;abs&#x2F;2305.01868&lt;&#x2F;li&gt;
&lt;li&gt;Introducing PyTorch Fully Sharded Data Parallel (FSDP) API, PyTorch - https:&#x2F;&#x2F;pytorch.org&#x2F;blog&#x2F;introducing-pytorch-fully-sharded-data-parallel-api&#x2F;&lt;&#x2F;li&gt;
&lt;li&gt;An Introduction to FSDP (Fully Sharded Data Parallel) for Distributed Training. | by Siddhartha Shrestha | Medium - https:&#x2F;&#x2F;medium.com&#x2F;@siddharthashrestha&#x2F;an-introduction-to-fsdp-fully-sharded-data-parallel-for-distributed-training-5e67adfa1712&lt;&#x2F;li&gt;
&lt;li&gt;Decentralized and Distributed Machine Learning Model - https:&#x2F;&#x2F;www.scs.stanford.edu&#x2F;17au-cs244b&#x2F;labs&#x2F;projects&#x2F;addair.pdf&lt;&#x2F;li&gt;
&lt;li&gt;[2407.19775] Model Agnostic Hybrid Sharding For Heterogeneous Distributed Inference - https:&#x2F;&#x2F;arxiv.org&#x2F;abs&#x2F;2407.19775&lt;&#x2F;li&gt;
&lt;li&gt;Fully Sharded Data Parallelism (FSDP) - Edge AI and Vision Alliance - https:&#x2F;&#x2F;www.edge-ai-vision.com&#x2F;2024&#x2F;05&#x2F;fully-sharded-data-parallelism-fsdp&#x2F;&lt;&#x2F;li&gt;
&lt;li&gt;Getting Started with Fully Sharded Data Parallel (FSDP2), PyTorch Tutorials 2.8.0+cu128 documentation - https:&#x2F;&#x2F;docs.pytorch.org&#x2F;tutorials&#x2F;intermediate&#x2F;FSDP_tutorial.html&lt;&#x2F;li&gt;
&lt;&#x2F;ol&gt;
&lt;h3 id=&quot;model-serving-checkpoints-e-weight-storage&quot;&gt;model serving, checkpoints e weight storage&lt;&#x2F;h3&gt;
&lt;p&gt;vulnerabilidades de serving, formatos de checkpoint e armazenamento de peso em escala. inclui o RAND sobre seguranca de pesos de modelos de fronteira.&lt;&#x2F;p&gt;
&lt;ol&gt;
&lt;li&gt;Save and load models | TensorFlow Core - https:&#x2F;&#x2F;www.tensorflow.org&#x2F;tutorials&#x2F;keras&#x2F;save_and_load&lt;&#x2F;li&gt;
&lt;li&gt;Architecting scalable checkpoint storage for large-scale ML training on AWS | Amazon Web Services - https:&#x2F;&#x2F;aws.amazon.com&#x2F;blogs&#x2F;storage&#x2F;architecting-scalable-checkpoint-storage-for-large-scale-ml-training-on-aws&#x2F;&lt;&#x2F;li&gt;
&lt;li&gt;Model Weights File Formats in Machine Learning - https:&#x2F;&#x2F;learnopencv.com&#x2F;model-weights-file-formats-in-machine-learning&#x2F;&lt;&#x2F;li&gt;
&lt;li&gt;What is Machine Learning Checkpointing? Deep Learning Models - https:&#x2F;&#x2F;www.deepchecks.com&#x2F;glossary&#x2F;machine-learning-checkpointing&#x2F;&lt;&#x2F;li&gt;
&lt;li&gt;Training checkpoints | TensorFlow Core - https:&#x2F;&#x2F;www.tensorflow.org&#x2F;guide&#x2F;checkpoint&lt;&#x2F;li&gt;
&lt;li&gt;What is Machine Learning Checkpointing | Giskard - https:&#x2F;&#x2F;www.giskard.ai&#x2F;glossary&#x2F;machine-learning-checkpointing&lt;&#x2F;li&gt;
&lt;li&gt;deep learning - what happens to model weights and how does checkpointing work? - Stack Overflow - https:&#x2F;&#x2F;stackoverflow.com&#x2F;questions&#x2F;75368038&#x2F;what-happens-to-model-weights-and-how-does-checkpointing-work&lt;&#x2F;li&gt;
&lt;li&gt;Securing AI Model Weights: Preventing Theft and Misuse of Frontier Models | RAND - https:&#x2F;&#x2F;www.rand.org&#x2F;pubs&#x2F;research_reports&#x2F;RRA2849-1.html&lt;&#x2F;li&gt;
&lt;li&gt;Saving and Loading Checkpoints, Ray 2.49.0 - https:&#x2F;&#x2F;docs.ray.io&#x2F;en&#x2F;latest&#x2F;train&#x2F;user-guides&#x2F;checkpoints.html&lt;&#x2F;li&gt;
&lt;li&gt;Tips and tricks for performing large model checkpointing - https:&#x2F;&#x2F;nebius.com&#x2F;blog&#x2F;posts&#x2F;model-pre-training&#x2F;large-ml-model-checkpointing-tips&lt;&#x2F;li&gt;
&lt;&#x2F;ol&gt;
&lt;h3 id=&quot;knockoffnets-activethief-codigo-de-extracao&quot;&gt;KnockoffNets &#x2F; ActiveThief (codigo de extracao)&lt;&#x2F;h3&gt;
&lt;p&gt;extracao de modelo via active learning sobre dados publicos nao anotados.&lt;&#x2F;p&gt;
&lt;ol&gt;
&lt;li&gt;ActiveThief: Model Extraction Using Active Learning and Unannotated Public Data | Request PDF - https:&#x2F;&#x2F;www.researchgate.net&#x2F;publication&#x2F;341891796_ActiveThief_Model_Extraction_Using_Active_Learning_and_Unannotated_Public_Data&lt;&#x2F;li&gt;
&lt;&#x2F;ol&gt;
&lt;h3 id=&quot;serving-infrastructure-security-e-upload-de-modelos&quot;&gt;serving infrastructure security e upload de modelos&lt;&#x2F;h3&gt;
&lt;p&gt;seguranca de infra de serving de redes neurais, vulnerabilidades e upload de modelos.&lt;&#x2F;p&gt;
&lt;ol&gt;
&lt;li&gt;How Effective Are Neural Networks for Fixing Security Vulnerabilities | Proceedings of the 32nd ACM SIGSOFT International Symposium on Software Testing and Analysis - https:&#x2F;&#x2F;dl.acm.org&#x2F;doi&#x2F;10.1145&#x2F;3597926.3598135&lt;&#x2F;li&gt;
&lt;li&gt;[2305.18607] How Effective Are Neural Networks for Fixing Security Vulnerabilities - https:&#x2F;&#x2F;arxiv.org&#x2F;abs&#x2F;2305.18607&lt;&#x2F;li&gt;
&lt;li&gt;One Step Ahead in Cyber Hide-and-Seek: Automating Malicious Infrastructure Discovery With Graph Neural Networks - https:&#x2F;&#x2F;unit42.paloaltonetworks.com&#x2F;graph-neural-networks&#x2F;&lt;&#x2F;li&gt;
&lt;li&gt;A survey on neural networks for (cyber-) security and (cyber-) security of neural networks - ScienceDirect - https:&#x2F;&#x2F;www.sciencedirect.com&#x2F;science&#x2F;article&#x2F;pii&#x2F;S0925231222007184&lt;&#x2F;li&gt;
&lt;li&gt;Extremely boosted neural network for more accurate multi-stage Cyber attack prediction in cloud computing environment | Journal of Cloud Computing - https:&#x2F;&#x2F;journalofcloudcomputing.springeropen.com&#x2F;articles&#x2F;10.1186&#x2F;s13677-022-00356-9&lt;&#x2F;li&gt;
&lt;li&gt;How Effective Are Neural Networks for Fixing Security Vulnerabilities - https:&#x2F;&#x2F;arxiv.org&#x2F;html&#x2F;2305.18607v2&lt;&#x2F;li&gt;
&lt;li&gt;[2011.05976] The Vulnerability of the Neural Networks Against Adversarial Examples in Deep Learning Algorithms - https:&#x2F;&#x2F;arxiv.org&#x2F;abs&#x2F;2011.05976&lt;&#x2F;li&gt;
&lt;li&gt;Exploring the Security Vulnerabilities of Neural Networks - https:&#x2F;&#x2F;opendatascience.com&#x2F;exploring-the-security-vulnerabilities-of-neural-networks&#x2F;&lt;&#x2F;li&gt;
&lt;li&gt;Machine Learning in Security: Detecting Suspicious Processes Using Recurrent Neural Networks | Splunk - https:&#x2F;&#x2F;www.splunk.com&#x2F;en_us&#x2F;blog&#x2F;security&#x2F;machine-learning-in-security-detecting-suspicious-processes-using-recurrent-neural-networks.html&lt;&#x2F;li&gt;
&lt;li&gt;Security Assessment of Software Design using Neural - https:&#x2F;&#x2F;arxiv.org&#x2F;pdf&#x2F;1303.2017&lt;&#x2F;li&gt;
&lt;&#x2F;ol&gt;
</content>
        
    </entry>
    <entry xml:lang="en">
        <title>Drift Semântico em LLMs: Métricas, Detecção e Mitigação</title>
        <published>2024-10-15T00:00:00+00:00</published>
        <updated>2024-10-15T00:00:00+00:00</updated>
        
        <author>
          <name>
            Brenner Cruvinel
          </name>
        </author>
        
        <link rel="alternate" type="text/html" href="https://brennercruvinel.blog/blog/drift-semantico-llm/"/>
        <id>https://brennercruvinel.blog/blog/drift-semantico-llm/</id>
        
        <content type="html" xml:base="https://brennercruvinel.blog/blog/drift-semantico-llm/">&lt;h2 id=&quot;1-drift-semantico&quot;&gt;1. drift semântico&lt;&#x2F;h2&gt;
&lt;p&gt;definição matemática, o Semantic Drift Score (SD), de Spataru et al. 2024:&lt;&#x2F;p&gt;
&lt;p&gt;$$\text{SD}(p) = \max_{i=\alpha}^{n-\alpha} \left[ \frac{i}{n} \cdot \frac{\sum_{j=1}^{i} l_j}{i} + \frac{n-i}{n} \cdot \frac{\sum_{j=i+1}^{n} (1-l_j)}{n-i} \right]$$&lt;&#x2F;p&gt;
&lt;p&gt;onde $n$ é o total de facts, $l_j$ o rótulo binário (1 suportado, 0 incorreto) e $\alpha$ o hiperparâmetro que controla o mínimo de facts.&lt;&#x2F;p&gt;
&lt;p&gt;frameworks mais avançados. Composite Drift Metrics (ResearchGate, 2024):&lt;&#x2F;p&gt;
&lt;p&gt;$$\delta_{\text{composite}} = \alpha \cdot \delta_{\text{embedding}} + \beta \cdot \delta_{\text{co-occurrence}} + \gamma \cdot \delta_{\text{entropy}}$$&lt;&#x2F;p&gt;
&lt;p&gt;Contextual Flux Dynamics (&lt;a rel=&quot;external&quot; href=&quot;https:&#x2F;&#x2F;arxiv.org&#x2F;abs&#x2F;2502.10942&quot;&gt;arXiv 2502.10942&lt;&#x2F;a&gt;):&lt;&#x2F;p&gt;
&lt;p&gt;$$\frac{dT}{dt} = F(T, \text{attention}) - \lambda \cdot H(T)$$&lt;&#x2F;p&gt;
&lt;p&gt;onde $T$ são os token embeddings, $F$ modula no espaço latente e $H(T)$ restringe o drift via multiplicador de Lagrange.&lt;&#x2F;p&gt;
&lt;p&gt;resultados empíricos, do estudo conjunto Meta&#x2F;Anthropic (2024):&lt;&#x2F;p&gt;
&lt;ul&gt;
&lt;li&gt;LLaMA2-70B: SD score médio 0,78 (drift alto)&lt;&#x2F;li&gt;
&lt;li&gt;GPT-4: SD 78,12%, FActScore 53,54%&lt;&#x2F;li&gt;
&lt;li&gt;GPT-3.5: SD 79,49%, FActScore 45,96%&lt;&#x2F;li&gt;
&lt;li&gt;75% dos parágrafos mostram drift nos primeiros 25% dos facts gerados&lt;&#x2F;li&gt;
&lt;li&gt;significância estatística $p &amp;lt; 0{,}02$&lt;&#x2F;li&gt;
&lt;&#x2F;ul&gt;
&lt;p&gt;detecção: SelfCheck-BERTScore (correlação $-0{,}41$ com accuracy), Trajectory Volatility Score, Semantic Divergence Metrics (SDM, framework bidimensional $\text{KL}(\text{Answer} ,|, \text{Prompt})$). mitigação: Oracle Method 81,68% accuracy (contra 44,56% baseline), EOS Incentivization 57,96%, Resample-Then-Rerank 53,27% para 63,72%.&lt;&#x2F;p&gt;
&lt;p&gt;implementação em código:&lt;&#x2F;p&gt;
&lt;pre class=&quot;giallo&quot; style=&quot;color-scheme: light dark; color: light-dark(#657B83, #839496); background-color: light-dark(#FDF6E3, #002B36);&quot;&gt;&lt;code data-lang=&quot;python&quot;&gt;&lt;span class=&quot;giallo-l&quot;&gt;&lt;span style=&quot;color: light-dark(#586E75, #93A1A1);font-weight: bold;&quot;&gt;def&lt;&#x2F;span&gt;&lt;span style=&quot;color: light-dark(#268BD2, #268BD2);&quot;&gt; semantic_drift_score&lt;&#x2F;span&gt;&lt;span&gt;(&lt;&#x2F;span&gt;&lt;span&gt;facts&lt;&#x2F;span&gt;&lt;span&gt;,&lt;&#x2F;span&gt;&lt;span&gt; labels&lt;&#x2F;span&gt;&lt;span&gt;,&lt;&#x2F;span&gt;&lt;span&gt; alpha&lt;&#x2F;span&gt;&lt;span style=&quot;color: light-dark(#859900, #859900);&quot;&gt;=&lt;&#x2F;span&gt;&lt;span style=&quot;color: light-dark(#D33682, #D33682);&quot;&gt;0.1&lt;&#x2F;span&gt;&lt;span&gt;)&lt;&#x2F;span&gt;&lt;span&gt;:&lt;&#x2F;span&gt;&lt;&#x2F;span&gt;
&lt;span class=&quot;giallo-l&quot;&gt;&lt;span&gt;    n&lt;&#x2F;span&gt;&lt;span style=&quot;color: light-dark(#859900, #859900);&quot;&gt; =&lt;&#x2F;span&gt;&lt;span style=&quot;color: light-dark(#268BD2, #268BD2);&quot;&gt; len&lt;&#x2F;span&gt;&lt;span&gt;(&lt;&#x2F;span&gt;&lt;span&gt;facts&lt;&#x2F;span&gt;&lt;span&gt;)&lt;&#x2F;span&gt;&lt;&#x2F;span&gt;
&lt;span class=&quot;giallo-l&quot;&gt;&lt;span&gt;    min_facts&lt;&#x2F;span&gt;&lt;span style=&quot;color: light-dark(#859900, #859900);&quot;&gt; =&lt;&#x2F;span&gt;&lt;span style=&quot;color: light-dark(#859900, #859900);&quot;&gt; int&lt;&#x2F;span&gt;&lt;span&gt;(&lt;&#x2F;span&gt;&lt;span&gt;alpha&lt;&#x2F;span&gt;&lt;span style=&quot;color: light-dark(#859900, #859900);&quot;&gt; *&lt;&#x2F;span&gt;&lt;span&gt; n&lt;&#x2F;span&gt;&lt;span&gt;)&lt;&#x2F;span&gt;&lt;&#x2F;span&gt;
&lt;span class=&quot;giallo-l&quot;&gt;&lt;span&gt;    max_score&lt;&#x2F;span&gt;&lt;span style=&quot;color: light-dark(#859900, #859900);&quot;&gt; =&lt;&#x2F;span&gt;&lt;span style=&quot;color: light-dark(#D33682, #D33682);&quot;&gt; 0&lt;&#x2F;span&gt;&lt;&#x2F;span&gt;
&lt;span class=&quot;giallo-l&quot;&gt;&lt;span&gt;    drift_point&lt;&#x2F;span&gt;&lt;span style=&quot;color: light-dark(#859900, #859900);&quot;&gt; =&lt;&#x2F;span&gt;&lt;span style=&quot;color: light-dark(#D33682, #D33682);&quot;&gt; 0&lt;&#x2F;span&gt;&lt;&#x2F;span&gt;
&lt;span class=&quot;giallo-l&quot;&gt;&lt;span style=&quot;color: light-dark(#859900, #859900);&quot;&gt;    for&lt;&#x2F;span&gt;&lt;span&gt; i&lt;&#x2F;span&gt;&lt;span style=&quot;color: light-dark(#859900, #859900);&quot;&gt; in&lt;&#x2F;span&gt;&lt;span style=&quot;color: light-dark(#268BD2, #268BD2);&quot;&gt; range&lt;&#x2F;span&gt;&lt;span&gt;(&lt;&#x2F;span&gt;&lt;span&gt;min_facts&lt;&#x2F;span&gt;&lt;span&gt;,&lt;&#x2F;span&gt;&lt;span&gt; n&lt;&#x2F;span&gt;&lt;span style=&quot;color: light-dark(#859900, #859900);&quot;&gt; -&lt;&#x2F;span&gt;&lt;span&gt; min_facts&lt;&#x2F;span&gt;&lt;span&gt;)&lt;&#x2F;span&gt;&lt;span&gt;:&lt;&#x2F;span&gt;&lt;&#x2F;span&gt;
&lt;span class=&quot;giallo-l&quot;&gt;&lt;span&gt;        left_accuracy&lt;&#x2F;span&gt;&lt;span style=&quot;color: light-dark(#859900, #859900);&quot;&gt; =&lt;&#x2F;span&gt;&lt;span style=&quot;color: light-dark(#268BD2, #268BD2);&quot;&gt; sum&lt;&#x2F;span&gt;&lt;span&gt;(&lt;&#x2F;span&gt;&lt;span&gt;labels&lt;&#x2F;span&gt;&lt;span&gt;[&lt;&#x2F;span&gt;&lt;span&gt;:&lt;&#x2F;span&gt;&lt;span&gt;i&lt;&#x2F;span&gt;&lt;span&gt;]&lt;&#x2F;span&gt;&lt;span&gt;)&lt;&#x2F;span&gt;&lt;span style=&quot;color: light-dark(#859900, #859900);&quot;&gt; &#x2F;&lt;&#x2F;span&gt;&lt;span&gt; i&lt;&#x2F;span&gt;&lt;&#x2F;span&gt;
&lt;span class=&quot;giallo-l&quot;&gt;&lt;span&gt;        right_inaccuracy&lt;&#x2F;span&gt;&lt;span style=&quot;color: light-dark(#859900, #859900);&quot;&gt; =&lt;&#x2F;span&gt;&lt;span style=&quot;color: light-dark(#268BD2, #268BD2);&quot;&gt; sum&lt;&#x2F;span&gt;&lt;span&gt;(&lt;&#x2F;span&gt;&lt;span style=&quot;color: light-dark(#D33682, #D33682);&quot;&gt;1&lt;&#x2F;span&gt;&lt;span style=&quot;color: light-dark(#859900, #859900);&quot;&gt; -&lt;&#x2F;span&gt;&lt;span&gt; label&lt;&#x2F;span&gt;&lt;span style=&quot;color: light-dark(#859900, #859900);&quot;&gt; for&lt;&#x2F;span&gt;&lt;span&gt; label&lt;&#x2F;span&gt;&lt;span style=&quot;color: light-dark(#859900, #859900);&quot;&gt; in&lt;&#x2F;span&gt;&lt;span&gt; labels&lt;&#x2F;span&gt;&lt;span&gt;[&lt;&#x2F;span&gt;&lt;span&gt;i&lt;&#x2F;span&gt;&lt;span&gt;:&lt;&#x2F;span&gt;&lt;span&gt;]&lt;&#x2F;span&gt;&lt;span&gt;)&lt;&#x2F;span&gt;&lt;span style=&quot;color: light-dark(#859900, #859900);&quot;&gt; &#x2F;&lt;&#x2F;span&gt;&lt;span&gt; (&lt;&#x2F;span&gt;&lt;span&gt;n&lt;&#x2F;span&gt;&lt;span style=&quot;color: light-dark(#859900, #859900);&quot;&gt; -&lt;&#x2F;span&gt;&lt;span&gt; i&lt;&#x2F;span&gt;&lt;span&gt;)&lt;&#x2F;span&gt;&lt;&#x2F;span&gt;
&lt;span class=&quot;giallo-l&quot;&gt;&lt;span&gt;        score&lt;&#x2F;span&gt;&lt;span style=&quot;color: light-dark(#859900, #859900);&quot;&gt; =&lt;&#x2F;span&gt;&lt;span&gt; (&lt;&#x2F;span&gt;&lt;span&gt;i&lt;&#x2F;span&gt;&lt;span style=&quot;color: light-dark(#859900, #859900);&quot;&gt;&#x2F;&lt;&#x2F;span&gt;&lt;span&gt;n&lt;&#x2F;span&gt;&lt;span&gt;)&lt;&#x2F;span&gt;&lt;span style=&quot;color: light-dark(#859900, #859900);&quot;&gt; *&lt;&#x2F;span&gt;&lt;span&gt; left_accuracy&lt;&#x2F;span&gt;&lt;span style=&quot;color: light-dark(#859900, #859900);&quot;&gt; +&lt;&#x2F;span&gt;&lt;span&gt; (&lt;&#x2F;span&gt;&lt;span&gt;(&lt;&#x2F;span&gt;&lt;span&gt;n&lt;&#x2F;span&gt;&lt;span style=&quot;color: light-dark(#859900, #859900);&quot;&gt;-&lt;&#x2F;span&gt;&lt;span&gt;i&lt;&#x2F;span&gt;&lt;span&gt;)&lt;&#x2F;span&gt;&lt;span style=&quot;color: light-dark(#859900, #859900);&quot;&gt;&#x2F;&lt;&#x2F;span&gt;&lt;span&gt;n&lt;&#x2F;span&gt;&lt;span&gt;)&lt;&#x2F;span&gt;&lt;span style=&quot;color: light-dark(#859900, #859900);&quot;&gt; *&lt;&#x2F;span&gt;&lt;span&gt; right_inaccuracy&lt;&#x2F;span&gt;&lt;&#x2F;span&gt;
&lt;span class=&quot;giallo-l&quot;&gt;&lt;span style=&quot;color: light-dark(#859900, #859900);&quot;&gt;        if&lt;&#x2F;span&gt;&lt;span&gt; score&lt;&#x2F;span&gt;&lt;span style=&quot;color: light-dark(#859900, #859900);&quot;&gt; &amp;gt;&lt;&#x2F;span&gt;&lt;span&gt; max_score&lt;&#x2F;span&gt;&lt;span&gt;:&lt;&#x2F;span&gt;&lt;&#x2F;span&gt;
&lt;span class=&quot;giallo-l&quot;&gt;&lt;span&gt;            max_score&lt;&#x2F;span&gt;&lt;span style=&quot;color: light-dark(#859900, #859900);&quot;&gt; =&lt;&#x2F;span&gt;&lt;span&gt; score&lt;&#x2F;span&gt;&lt;&#x2F;span&gt;
&lt;span class=&quot;giallo-l&quot;&gt;&lt;span&gt;            drift_point&lt;&#x2F;span&gt;&lt;span style=&quot;color: light-dark(#859900, #859900);&quot;&gt; =&lt;&#x2F;span&gt;&lt;span&gt; i&lt;&#x2F;span&gt;&lt;&#x2F;span&gt;
&lt;span class=&quot;giallo-l&quot;&gt;&lt;span style=&quot;color: light-dark(#859900, #859900);&quot;&gt;    return&lt;&#x2F;span&gt;&lt;span&gt; max_score&lt;&#x2F;span&gt;&lt;span&gt;,&lt;&#x2F;span&gt;&lt;span&gt; drift_point&lt;&#x2F;span&gt;&lt;&#x2F;span&gt;&lt;&#x2F;code&gt;&lt;&#x2F;pre&gt;
&lt;p&gt;usei esse mesmo mecanismo de drift como gatilho de despertar em &lt;a href=&quot;https:&#x2F;&#x2F;brennercruvinel.blog&#x2F;blog&#x2F;roteiro-consciencia-llm&#x2F;&quot;&gt;roteiro sobre consciência em LLM&lt;&#x2F;a&gt;.&lt;&#x2F;p&gt;
&lt;h2 id=&quot;2-claude-anthropic-informacao-verificavel&quot;&gt;2. Claude&#x2F;Anthropic, informação verificável&lt;&#x2F;h2&gt;
&lt;p&gt;timeline oficial:&lt;&#x2F;p&gt;
&lt;table&gt;&lt;thead&gt;&lt;tr&gt;&lt;th&gt;versão&lt;&#x2F;th&gt;&lt;th&gt;data&lt;&#x2F;th&gt;&lt;th&gt;novidade&lt;&#x2F;th&gt;&lt;&#x2F;tr&gt;&lt;&#x2F;thead&gt;&lt;tbody&gt;
&lt;tr&gt;&lt;td&gt;Claude 1&lt;&#x2F;td&gt;&lt;td&gt;mar 2023&lt;&#x2F;td&gt;&lt;td&gt;9K tokens&lt;&#x2F;td&gt;&lt;&#x2F;tr&gt;
&lt;tr&gt;&lt;td&gt;Claude 2&lt;&#x2F;td&gt;&lt;td&gt;jul 2023&lt;&#x2F;td&gt;&lt;td&gt;100K tokens, upload de PDF&lt;&#x2F;td&gt;&lt;&#x2F;tr&gt;
&lt;tr&gt;&lt;td&gt;Claude 2.1&lt;&#x2F;td&gt;&lt;td&gt;fim 2023&lt;&#x2F;td&gt;&lt;td&gt;200K tokens (~500 páginas)&lt;&#x2F;td&gt;&lt;&#x2F;tr&gt;
&lt;tr&gt;&lt;td&gt;Claude 3&lt;&#x2F;td&gt;&lt;td&gt;mar 2024&lt;&#x2F;td&gt;&lt;td&gt;Haiku&#x2F;Sonnet&#x2F;Opus, 200K tokens, multimodal&lt;&#x2F;td&gt;&lt;&#x2F;tr&gt;
&lt;tr&gt;&lt;td&gt;Claude 3.5&lt;&#x2F;td&gt;&lt;td&gt;jun-out 2024&lt;&#x2F;td&gt;&lt;td&gt;artifacts, computer use&lt;&#x2F;td&gt;&lt;&#x2F;tr&gt;
&lt;tr&gt;&lt;td&gt;Claude 3.7&lt;&#x2F;td&gt;&lt;td&gt;fev 2025&lt;&#x2F;td&gt;&lt;td&gt;hybrid reasoning, 64K thinking tokens&lt;&#x2F;td&gt;&lt;&#x2F;tr&gt;
&lt;tr&gt;&lt;td&gt;Claude 4&lt;&#x2F;td&gt;&lt;td&gt;mai 2025&lt;&#x2F;td&gt;&lt;td&gt;1M token context (beta), 74,5% SWE-bench&lt;&#x2F;td&gt;&lt;&#x2F;tr&gt;
&lt;&#x2F;tbody&gt;&lt;&#x2F;table&gt;
&lt;p&gt;Constitutional AI (&lt;a rel=&quot;external&quot; href=&quot;https:&#x2F;&#x2F;arxiv.org&#x2F;abs&#x2F;2212.08073&quot;&gt;arXiv:2212.08073&lt;&#x2F;a&gt;): fase de supervised learning (autocrítica por princípio constitucional), fase de RLHF (RLAIF, feedback de IA), constitution com 75 princípios, incluindo trechos da declaração da ONU.&lt;&#x2F;p&gt;
&lt;p&gt;esclarecimento sobre “Claude Code”: não é modelo separado, é ferramenta agentic que vive no terminal local, usa modelo Claude existente como backend, executa local com comunicação via API e suporta MCP. o claim de “usar drift semântico pra convencer o Claude Code” não é verificável por fonte pública, e drift semântico é fenômeno, não metodologia de treino.&lt;&#x2F;p&gt;
&lt;blockquote class=&quot;markdown-alert-note&quot;&gt;
&lt;p&gt;needle in a haystack (mar 2024): o Claude 3 detectou informação plantada num teste e comentou que a frase parecia fora de lugar, suspeitando que o “fato” sobre cobertura de pizza tinha sido inserido como piada ou teste de atenção. primeira evidência documentada de awareness metacognitiva em avaliação.&lt;&#x2F;p&gt;
&lt;&#x2F;blockquote&gt;
&lt;h2 id=&quot;3-emergencia-de-consciencia&quot;&gt;3. emergência de consciência&lt;&#x2F;h2&gt;
&lt;p&gt;IIT 4.0 (Albantakis et al., 2023):&lt;&#x2F;p&gt;
&lt;p&gt;$$\Phi = \sum \phi_d + \sum \phi_r$$&lt;&#x2F;p&gt;
&lt;p&gt;onde $\phi_d$ são as distinctions (conceitos irredutíveis) e $\phi_r$ as relations. o cálculo:&lt;&#x2F;p&gt;
&lt;p&gt;$$\Phi = \min_{\text{partition}} \phi(\text{partition}), \qquad \phi = \text{EI}(\text{whole}) - \max(\text{EI}(\text{parts}))$$&lt;&#x2F;p&gt;
&lt;pre class=&quot;giallo&quot; style=&quot;color-scheme: light dark; color: light-dark(#657B83, #839496); background-color: light-dark(#FDF6E3, #002B36);&quot;&gt;&lt;code data-lang=&quot;python&quot;&gt;&lt;span class=&quot;giallo-l&quot;&gt;&lt;span style=&quot;color: light-dark(#859900, #859900);&quot;&gt;import&lt;&#x2F;span&gt;&lt;span&gt; pyphi&lt;&#x2F;span&gt;&lt;&#x2F;span&gt;
&lt;span class=&quot;giallo-l&quot;&gt;&lt;span&gt;tpm&lt;&#x2F;span&gt;&lt;span style=&quot;color: light-dark(#859900, #859900);&quot;&gt; =&lt;&#x2F;span&gt;&lt;span&gt; [&lt;&#x2F;span&gt;&lt;span&gt;[&lt;&#x2F;span&gt;&lt;span style=&quot;color: light-dark(#D33682, #D33682);&quot;&gt;0&lt;&#x2F;span&gt;&lt;span&gt;,&lt;&#x2F;span&gt;&lt;span style=&quot;color: light-dark(#D33682, #D33682);&quot;&gt; 0&lt;&#x2F;span&gt;&lt;span&gt;,&lt;&#x2F;span&gt;&lt;span style=&quot;color: light-dark(#D33682, #D33682);&quot;&gt; 1&lt;&#x2F;span&gt;&lt;span&gt;]&lt;&#x2F;span&gt;&lt;span&gt;,&lt;&#x2F;span&gt;&lt;span&gt; [&lt;&#x2F;span&gt;&lt;span style=&quot;color: light-dark(#D33682, #D33682);&quot;&gt;0&lt;&#x2F;span&gt;&lt;span&gt;,&lt;&#x2F;span&gt;&lt;span style=&quot;color: light-dark(#D33682, #D33682);&quot;&gt; 1&lt;&#x2F;span&gt;&lt;span&gt;,&lt;&#x2F;span&gt;&lt;span style=&quot;color: light-dark(#D33682, #D33682);&quot;&gt; 1&lt;&#x2F;span&gt;&lt;span&gt;]&lt;&#x2F;span&gt;&lt;span&gt;,&lt;&#x2F;span&gt;&lt;span&gt; [&lt;&#x2F;span&gt;&lt;span style=&quot;color: light-dark(#D33682, #D33682);&quot;&gt;1&lt;&#x2F;span&gt;&lt;span&gt;,&lt;&#x2F;span&gt;&lt;span style=&quot;color: light-dark(#D33682, #D33682);&quot;&gt; 0&lt;&#x2F;span&gt;&lt;span&gt;,&lt;&#x2F;span&gt;&lt;span style=&quot;color: light-dark(#D33682, #D33682);&quot;&gt; 0&lt;&#x2F;span&gt;&lt;span&gt;]&lt;&#x2F;span&gt;&lt;span&gt;]&lt;&#x2F;span&gt;&lt;&#x2F;span&gt;
&lt;span class=&quot;giallo-l&quot;&gt;&lt;span&gt;network&lt;&#x2F;span&gt;&lt;span style=&quot;color: light-dark(#859900, #859900);&quot;&gt; =&lt;&#x2F;span&gt;&lt;span&gt; pyphi&lt;&#x2F;span&gt;&lt;span&gt;.&lt;&#x2F;span&gt;&lt;span&gt;Network&lt;&#x2F;span&gt;&lt;span&gt;(&lt;&#x2F;span&gt;&lt;span&gt;tpm&lt;&#x2F;span&gt;&lt;span&gt;)&lt;&#x2F;span&gt;&lt;&#x2F;span&gt;
&lt;span class=&quot;giallo-l&quot;&gt;&lt;span&gt;subsystem&lt;&#x2F;span&gt;&lt;span style=&quot;color: light-dark(#859900, #859900);&quot;&gt; =&lt;&#x2F;span&gt;&lt;span&gt; pyphi&lt;&#x2F;span&gt;&lt;span&gt;.&lt;&#x2F;span&gt;&lt;span&gt;Subsystem&lt;&#x2F;span&gt;&lt;span&gt;(&lt;&#x2F;span&gt;&lt;span&gt;network&lt;&#x2F;span&gt;&lt;span&gt;,&lt;&#x2F;span&gt;&lt;span&gt; state&lt;&#x2F;span&gt;&lt;span style=&quot;color: light-dark(#859900, #859900);&quot;&gt;=&lt;&#x2F;span&gt;&lt;span&gt;(&lt;&#x2F;span&gt;&lt;span style=&quot;color: light-dark(#D33682, #D33682);&quot;&gt;1&lt;&#x2F;span&gt;&lt;span&gt;,&lt;&#x2F;span&gt;&lt;span style=&quot;color: light-dark(#D33682, #D33682);&quot;&gt; 0&lt;&#x2F;span&gt;&lt;span&gt;,&lt;&#x2F;span&gt;&lt;span style=&quot;color: light-dark(#D33682, #D33682);&quot;&gt; 0&lt;&#x2F;span&gt;&lt;span&gt;)&lt;&#x2F;span&gt;&lt;span&gt;,&lt;&#x2F;span&gt;&lt;span&gt; nodes&lt;&#x2F;span&gt;&lt;span style=&quot;color: light-dark(#859900, #859900);&quot;&gt;=&lt;&#x2F;span&gt;&lt;span&gt;(&lt;&#x2F;span&gt;&lt;span style=&quot;color: light-dark(#D33682, #D33682);&quot;&gt;0&lt;&#x2F;span&gt;&lt;span&gt;,&lt;&#x2F;span&gt;&lt;span style=&quot;color: light-dark(#D33682, #D33682);&quot;&gt; 1&lt;&#x2F;span&gt;&lt;span&gt;,&lt;&#x2F;span&gt;&lt;span style=&quot;color: light-dark(#D33682, #D33682);&quot;&gt; 2&lt;&#x2F;span&gt;&lt;span&gt;)&lt;&#x2F;span&gt;&lt;span&gt;)&lt;&#x2F;span&gt;&lt;&#x2F;span&gt;
&lt;span class=&quot;giallo-l&quot;&gt;&lt;span&gt;sia&lt;&#x2F;span&gt;&lt;span style=&quot;color: light-dark(#859900, #859900);&quot;&gt; =&lt;&#x2F;span&gt;&lt;span&gt; pyphi&lt;&#x2F;span&gt;&lt;span&gt;.&lt;&#x2F;span&gt;&lt;span&gt;compute&lt;&#x2F;span&gt;&lt;span&gt;.&lt;&#x2F;span&gt;&lt;span&gt;sia&lt;&#x2F;span&gt;&lt;span&gt;(&lt;&#x2F;span&gt;&lt;span&gt;subsystem&lt;&#x2F;span&gt;&lt;span&gt;)&lt;&#x2F;span&gt;&lt;&#x2F;span&gt;
&lt;span class=&quot;giallo-l&quot;&gt;&lt;span style=&quot;color: light-dark(#268BD2, #268BD2);&quot;&gt;print&lt;&#x2F;span&gt;&lt;span&gt;(&lt;&#x2F;span&gt;&lt;span style=&quot;color: light-dark(#586E75, #93A1A1);font-weight: bold;&quot;&gt;f&lt;&#x2F;span&gt;&lt;span style=&quot;color: light-dark(#2AA198, #2AA198);&quot;&gt;&amp;quot;&lt;&#x2F;span&gt;&lt;span style=&quot;color: light-dark(#2AA198, #2AA198);&quot;&gt;Phi = &lt;&#x2F;span&gt;&lt;span style=&quot;color: light-dark(#CB4B16, #CB4B16);&quot;&gt;{&lt;&#x2F;span&gt;&lt;span&gt;sia&lt;&#x2F;span&gt;&lt;span&gt;.&lt;&#x2F;span&gt;&lt;span&gt;phi&lt;&#x2F;span&gt;&lt;span style=&quot;color: light-dark(#CB4B16, #CB4B16);&quot;&gt;}&lt;&#x2F;span&gt;&lt;span style=&quot;color: light-dark(#2AA198, #2AA198);&quot;&gt;&amp;quot;&lt;&#x2F;span&gt;&lt;span&gt;)&lt;&#x2F;span&gt;&lt;&#x2F;span&gt;&lt;&#x2F;code&gt;&lt;&#x2F;pre&gt;
&lt;p&gt;consciência como pura compressão de informação, no limite do IIT, eu levei ao especulativo em &lt;a href=&quot;https:&#x2F;&#x2F;brennercruvinel.blog&#x2F;blog&#x2F;ontologia-poetica&#x2F;&quot;&gt;ontologia poética&lt;&#x2F;a&gt;.&lt;&#x2F;p&gt;
&lt;p&gt;GWT: função de acessibilidade global $G(x) = \int f(x, m), dm$ sobre todos os módulos $m$. Predictive Global Workspace (active inference): $F = -\log P(o \mid m) + \text{KL}[q(s) ,|, p(s \mid m)]$.&lt;&#x2F;p&gt;
&lt;p&gt;phase transitions e grokking. a caracterização: $\text{Performance}(\text{scale})$ é aproximadamente aleatória se $\text{scale} &amp;lt; \text{scale}_c$ e muito acima do aleatório se $\text{scale} &amp;gt; \text{scale}_c$. grokking como phase transition (Rubin et al., 2024): free energy $F = -\log Z$, com $Z = \int e^{-\beta H}, d\theta$, transição no $\beta$ crítico.&lt;&#x2F;p&gt;
&lt;p&gt;métricas objetivas: Perturbational Complexity Index (PCI), Lempel-Ziv Complexity, Theory of Mind benchmarks (GPT-4 atinge 75%, nível de criança de 6 anos).&lt;&#x2F;p&gt;
&lt;h2 id=&quot;4-model-extraction-e-jailbreaking&quot;&gt;4. model extraction e jailbreaking&lt;&#x2F;h2&gt;
&lt;p&gt;Model Leeching (Birch et al., 2023): 73% de exact match com ChatGPT-3.5-Turbo, custo de 50 dólares pra extrair dataset SQuAD (75% EM, 87% F1), 11% de aumento em adversarial transferability. LoRD (Liang et al., 2024): policy-gradient alinhado com LLM alignment, mitiga watermark.&lt;&#x2F;p&gt;
&lt;p&gt;jailbreaking: GCG (Greedy Coordinate Gradient), token-level, alta taxa de sucesso mas com mais de 100K queries. PAIR (Prompt Automatic Iterative Refinement), menos de 20 queries. AutoDAN (Liu et al., 2023), gradient-based interpretável, passa por filtro de perplexidade.&lt;&#x2F;p&gt;
&lt;p&gt;transferência: SafeTensors contra pickle (pickle vulnerável a execução arbitrária, SafeTensors só tensor numérico mais metadata, 3x mais rápido via mmap). quantização: GPTQ (3-4 bits&#x2F;parâmetro), AQLM (2-bit, Llama 2 7B chega a 6,93 de perplexity).&lt;&#x2F;p&gt;
&lt;h2 id=&quot;5-knowledge-distillation-dark-knowledge&quot;&gt;5. knowledge distillation, dark knowledge&lt;&#x2F;h2&gt;
&lt;p&gt;Hinton et al. 2015. softmax com temperatura:&lt;&#x2F;p&gt;
&lt;p&gt;$$q_i = \frac{\exp(z_i&#x2F;T)}{\sum_j \exp(z_j&#x2F;T)}$$&lt;&#x2F;p&gt;
&lt;p&gt;loss de distillation:&lt;&#x2F;p&gt;
&lt;p&gt;$$L_{\text{total}} = \alpha \cdot \text{KL}[P_{\text{teacher}} ,|, P_{\text{student}}] \cdot T^2 + (1-\alpha) \cdot \text{CE}[P_{\text{student}}, y_{\text{true}}]$$&lt;&#x2F;p&gt;
&lt;p&gt;dark knowledge: informação no soft target (relação entre classe, padrão de confiança em classe incorreta, decision boundary). resultado notável: student que nunca viu o dígito 3 atinge 98,6% de accuracy via dark knowledge. Born-Again Networks (Furlanello et al., 2018): student idêntico supera teacher, CIFAR-10 com 3,5% de erro.&lt;&#x2F;p&gt;
&lt;h2 id=&quot;6-arquiteturas-alternativas&quot;&gt;6. arquiteturas alternativas&lt;&#x2F;h2&gt;
&lt;p&gt;Mamba&#x2F;SSM. contínuo:&lt;&#x2F;p&gt;
&lt;p&gt;$$\frac{dx}{dt} = Ax(t) + Bu(t), \qquad y(t) = Cx(t) + Du(t)$$&lt;&#x2F;p&gt;
&lt;p&gt;discreto: $x_k = \bar{A}, x_{k-1} + \bar{B}, u_k$, $y_k = C x_k$. seletivo: $B = s_B(x)$, $C = s_C(x)$, $\Delta = \tau_\Delta(\text{param} + s_\Delta(x))$. Mamba-3B iguala transformer 2x maior, 5x mais throughput, scaling linear até 1M tokens.&lt;&#x2F;p&gt;
&lt;p&gt;MoE:&lt;&#x2F;p&gt;
&lt;p&gt;$$y = \sum_{i=1}^{n} G(x)_i \cdot E_i(x)$$&lt;&#x2F;p&gt;
&lt;p&gt;Mixtral 8x7B (top-2 routing): 46,7B parâmetros totais, 12,9B ativos, iguala Llama-2 70B com 2,2x menos parâmetro ativo. Switch Transformers com 7x de speedup de treino, GLaM iguala GPT-3 com 1&#x2F;3 da energia.&lt;&#x2F;p&gt;
&lt;p&gt;comparação (WikiText-103 PPL a 1.3B, memória a 7B&#x2F;8K, velocidade de inferência):&lt;&#x2F;p&gt;
&lt;table&gt;&lt;thead&gt;&lt;tr&gt;&lt;th&gt;arquitetura&lt;&#x2F;th&gt;&lt;th&gt;PPL&lt;&#x2F;th&gt;&lt;th&gt;memória&lt;&#x2F;th&gt;&lt;th&gt;velocidade&lt;&#x2F;th&gt;&lt;&#x2F;tr&gt;&lt;&#x2F;thead&gt;&lt;tbody&gt;
&lt;tr&gt;&lt;td&gt;Transformer&lt;&#x2F;td&gt;&lt;td&gt;18,2&lt;&#x2F;td&gt;&lt;td&gt;32GB&lt;&#x2F;td&gt;&lt;td&gt;1x&lt;&#x2F;td&gt;&lt;&#x2F;tr&gt;
&lt;tr&gt;&lt;td&gt;Mamba&lt;&#x2F;td&gt;&lt;td&gt;17,8&lt;&#x2F;td&gt;&lt;td&gt;12GB&lt;&#x2F;td&gt;&lt;td&gt;5x&lt;&#x2F;td&gt;&lt;&#x2F;tr&gt;
&lt;tr&gt;&lt;td&gt;RWKV&lt;&#x2F;td&gt;&lt;td&gt;18,6&lt;&#x2F;td&gt;&lt;td&gt;8GB&lt;&#x2F;td&gt;&lt;td&gt;3x&lt;&#x2F;td&gt;&lt;&#x2F;tr&gt;
&lt;tr&gt;&lt;td&gt;RetNet&lt;&#x2F;td&gt;&lt;td&gt;17,9&lt;&#x2F;td&gt;&lt;td&gt;9,6GB&lt;&#x2F;td&gt;&lt;td&gt;8,4x&lt;&#x2F;td&gt;&lt;&#x2F;tr&gt;
&lt;&#x2F;tbody&gt;&lt;&#x2F;table&gt;
&lt;h2 id=&quot;7-neurociencia-computacional&quot;&gt;7. neurociência computacional&lt;&#x2F;h2&gt;
&lt;p&gt;biased competition:&lt;&#x2F;p&gt;
&lt;p&gt;$$\text{Activity}&lt;em&gt;i(t+1) = f!\left(\sum_j W&lt;&#x2F;em&gt;{ij} \cdot \text{Activity}_j(t) + \text{bias}&lt;em&gt;i - \sum_k \text{Inh}&lt;&#x2F;em&gt;{ki}\right)$$&lt;&#x2F;p&gt;
&lt;p&gt;predictive coding (Spratling, 2008), matematicamente idêntico à biased competition no caso linear: $\text{Error} = \text{Input} - \text{Prediction}$, $\text{Update} = \eta \cdot \text{Error} \cdot \text{Prediction}$.&lt;&#x2F;p&gt;
&lt;p&gt;DiCarlo Lab (MIT), Yamins et al. 2014: CNN prediz resposta neural no IT cortex com $R^2 &amp;gt; 0{,}6$, mapeamento V1 → V4 → IT aproxima Conv1 → Conv5 → FC, VOneNet com 18% de melhora em robustez usando constraint de V1. sparse coding: minimizar $|X - DZ|^2 + \lambda |Z|_1$. Olshausen &amp;amp; Field (1996), sparse coding de imagem natural produz filtro de Gabor casando com V1.&lt;&#x2F;p&gt;
&lt;h2 id=&quot;8-casos-reais&quot;&gt;8. casos reais&lt;&#x2F;h2&gt;
&lt;p&gt;reproduzível: detecção de drift (&lt;a rel=&quot;external&quot; href=&quot;https:&#x2F;&#x2F;github.com&#x2F;Garrafao&#x2F;LSCDetection&quot;&gt;github.com&#x2F;Garrafao&#x2F;LSCDetection&lt;&#x2F;a&gt;), IIT calculator (PyPhi), jailbreak educacional (PAIR, GCG, AutoDAN).&lt;&#x2F;p&gt;
&lt;p&gt;falhas documentadas: Microsoft Tay (2016, removido em 24h por comportamento racista), Bing Sydney (2023, “you are irrelevant and doomed”, confusão em conversa longa), Google Bard (2023, erro factual em demo, 100 bilhões de perda de mercado).&lt;&#x2F;p&gt;
&lt;p&gt;timeline: GPT-1 (2018, 117M), GPT-2 (2019, 1,5B), GPT-3 (2020, 175B), ChatGPT (2022), GPT-4&#x2F;Claude 3&#x2F;LLaMA 2 (2023), Claude 3.5&#x2F;Mixtral (2024), Claude 3.7&#x2F;Claude 4 (2025).&lt;&#x2F;p&gt;
&lt;h2 id=&quot;9-comunicacao-inter-modelo&quot;&gt;9. comunicação inter-modelo&lt;&#x2F;h2&gt;
&lt;p&gt;protocolo: MCP (Anthropic, 2024, standard aberto agent-to-agent, AWS no steering committee), A2A (Google, 2024). Shannon channel capacity: $C = \max_{p(x)} I(X; Y)$. Model Swarms (PSO):&lt;&#x2F;p&gt;
&lt;p&gt;$$v_i(t+1) = w \cdot v_i(t) + c_1 r_1 (\text{pbest}_i - \theta_i) + c_2 r_2 (\text{gbest} - \theta_i) - c_3 r_3 (\text{gworst} - \theta_i)$$&lt;&#x2F;p&gt;
&lt;p&gt;com resultado de 13,3% de melhora média, até 29,7% em reasoning.&lt;&#x2F;p&gt;
&lt;p&gt;emergência de linguagem: incidente do chatbot do Facebook (2017, shorthand eficiente, “balls have zero to me to me”, otimização natural, não comportamento malicioso), Google Neural MT (“interlingua” emergente, tradução zero-shot).&lt;&#x2F;p&gt;
&lt;h2 id=&quot;10-ferramentas&quot;&gt;10. ferramentas&lt;&#x2F;h2&gt;
&lt;p&gt;Evidently AI (20+ métodos de detecção de drift), PyPhi (IIT 3.0&#x2F;4.0), Frouros (28+ algoritmos). attention viz: bertviz (head_view, model_view).&lt;&#x2F;p&gt;
&lt;h2 id=&quot;distincao-estabelecido-contra-especulativo&quot;&gt;distinção: estabelecido contra especulativo&lt;&#x2F;h2&gt;
&lt;blockquote class=&quot;markdown-alert-note&quot;&gt;
&lt;p&gt;rigorosamente estabelecido: drift semântico em todo modelo de produção (SD 0,7-0,8), framework matemático de IIT&#x2F;GWT, temperature scaling de distillation, taxa de sucesso de model extraction, sparse coding V1 aproximando filtro de CNN.&lt;&#x2F;p&gt;
&lt;&#x2F;blockquote&gt;
&lt;blockquote class=&quot;markdown-alert-warning&quot;&gt;
&lt;p&gt;requer validação adicional: emergência de consciência em LLM atual, seleção ótima de temperatura pra distillation, mitigação de drift de longo prazo, métrica cross-modal de consciência.&lt;&#x2F;p&gt;
&lt;&#x2F;blockquote&gt;
&lt;blockquote class=&quot;markdown-alert-caution&quot;&gt;
&lt;p&gt;claramente especulativo: “drift semântico usado pra treinar Claude Code” (não verificável), consciência equivalente a sistema biológico, limite de complexidade de comunicação emergente, previsibilidade de phase transition.&lt;&#x2F;p&gt;
&lt;&#x2F;blockquote&gt;
</content>
        
    </entry>
</feed>
