Model Extraction Attacks contra Redes Neurais Black-Box (2020-2026)

By Brenner Cruvinel • 8 minutes read

Model Extraction Attacks on Black-Box Neural Networks (2020-2026)

monorep fresh co, ataques de extracao contra modelos black-box, surveys e literatura agregada.

  1. B3: Backdoor Attacks against Black-box Machine Learning Models | ACM Transactions on Privacy and Security - https://dl.acm.org/doi/10.1145/3605212
  2. [PDF] CloudLeak: Large-Scale Deep Learning Models Stealing Through Adversarial Examples | Semantic Scholar - https://www.semanticscholar.org/paper/CloudLeak:-Large-Scale-Deep-Learning-Models-Through-Yu-Yang/4d548fd21aad60e3052455e22b7a57cc1f06e3c3
  3. Stealing Neural Networks With Model Extraction Attacks, Dr. Nicholas Carlini | Commonwealth Cyber Initiative (CCI) | Virginia Tech - https://cyberinitiative.org/events-programs/2020/stealing-neural-networks-with-model-extraction-attacks-dr-nich.html
  4. I Know What You Trained Last Summer: A Survey on Stealing Machine Learning Models and Defences | ACM Computing Surveys - https://dl.acm.org/doi/10.1145/3595292
  5. GitHub - kzhao5/Model-Extraction-Stealing-Attacks-Machine-Learning-Literature - https://github.com/kzhao5/Model-Extraction-Stealing-Attacks-Machine-Learning-Literature
  6. First Model-Stealing Attack Reveals Secrets of Black-Box Production Language Models | Synced - https://syncedreview.com/2024/03/27/first-model-stealing-attack-reveals-secrets-of-black-box-production-language-models/
  7. A realistic model extraction attack against graph neural networks - ScienceDirect - https://www.sciencedirect.com/science/article/abs/pii/S0950705124007780
  8. [1602.02697] Practical Black-Box Attacks against Machine Learning - https://arxiv.org/abs/1602.02697
  9. Practical Black-Box Attacks against Machine Learning | Proceedings of the 2017 ACM on Asia Conference on Computer and Communications Security - https://dl.acm.org/doi/10.1145/3052973.3053009
  10. Black-box targeted adversarial attacks for deep neural networks - ScienceDirect - https://www.sciencedirect.com/science/article/abs/pii/S0925231225016893

Stealing Machine Learning Models via Prediction APIs (arxiv)

o paper seminal do Tramer e suas variantes, mais codigo de extracao via prediction API.

  1. [1609.02943] Stealing Machine Learning Models via Prediction APIs - https://arxiv.org/abs/1609.02943
  2. Stealing Machine Learning Models via Prediction APIs - https://www.usenix.org/system/files/conference/usenixsecurity16/sec16_paper_tramer.pdf
  3. Stealing machine learning models via prediction APIs | Proceedings of the 25th USENIX Conference on Security Symposium - https://dl.acm.org/doi/10.5555/3241094.3241142
  4. [1609.02943] Stealing Machine Learning Models via Prediction APIs (HTML version) - https://ar5iv.labs.arxiv.org/html/1609.02943
  5. [1609.02943v2] Stealing Machine Learning Models via Prediction APIs (Updated version) - https://arxiv.org/abs/1609.02943v2
  6. Efficient and Effective Model Extraction - https://arxiv.org/html/2409.14122v1
  7. Scholars@Duke publication: Stealing Machine Learning Models via Prediction APIs - https://scholars.duke.edu/publication/1493919
  8. GitHub - ftramer/Steal-ML: Model extraction attacks on Machine-Learning-as-a-Service platforms - https://github.com/ftramer/Steal-ML
  9. [PDF] Stealing Machine Learning Models via Prediction APIs | Semantic Scholar - https://www.semanticscholar.org/paper/Stealing-Machine-Learning-Models-via-Prediction-Tram%C3%A8r-Zhang/d15b21bdd117877c2d0e865b17d6a336737aea99
  10. Stealing Machine Learning Models via Prediction APIs | USENIX - https://www.usenix.org/conference/usenixsecurity16/technical-sessions/presentation/tramer

model inversion attacks 2020-2025 (deep leakage from gradients)

inversao de modelo e vazamento por gradiente, com foco em federated learning e training de language models.

  1. Iterative and mixed-spaces image gradient inversion attack in federated learning | Cybersecurity - https://cybersecurity.springeropen.com/articles/10.1186/s42400-024-00227-7
  2. Understanding Deep Gradient Leakage via Inversion Influence Functions - PubMed - https://pubmed.ncbi.nlm.nih.gov/38606303/
  3. Deep leakage from gradients | Proceedings of the 33rd International Conference on Neural Information Processing Systems - https://dl.acm.org/doi/10.5555/3454287.3455610
  4. Wasserstein Distance-Based Deep Leakage from Gradients - PMC - https://www.ncbi.nlm.nih.gov/pmc/articles/PMC10217429/
  5. Uncovering Gradient Inversion Risks in Practical Language Model Training | Proceedings of the 2024 on ACM SIGSAC Conference on Computer and Communications Security - https://dl.acm.org/doi/10.1145/3658644.3690292
  6. Deep Leakage from Gradients | SpringerLink - https://link.springer.com/chapter/10.1007/978-3-030-63076-8_2
  7. Understanding Deep Gradient Leakage via Inversion Influence Functions - https://arxiv.org/html/2309.13016
  8. [1906.08935] Deep Leakage from Gradients - https://arxiv.org/abs/1906.08935
  9. Inverting Gradient Attacks Naturally Makes Data Poisons: An Availability Attack on Neural Networks - https://arxiv.org/html/2410.21453v1
  10. A Survey of the Implementations of Model Inversion Attacks | SpringerLink - https://link.springer.com/chapter/10.1007/978-3-031-30648-8_1

ONNX, SafeTensors, serializacao de pesos

formatos de armazenamento de peso de rede neural e conversao entre eles. relevante pra entender superficie de ataque na serializacao.

  1. ONNX - https://huggingface.co/docs/transformers/en/serialization
  2. Safetensors, CKPT, ONNX, GGUF, and Other Key AI Model Formats | NEDNEX - https://nednex.com/en/what-are-safetensors/
  3. Model Weights File Formats in Machine Learning - https://learnopencv.com/model-weights-file-formats-in-machine-learning/
  4. Export to ONNX - https://huggingface.co/docs/transformers/v4.36.1/serialization
  5. justinchuby/onnx-safetensors | DeepWiki - https://deepwiki.com/justinchuby/onnx-safetensors
  6. Common AI Model Formats - https://huggingface.co/blog/ngxson/common-ai-model-formats
  7. python - How to convert safetensors model to onnx model? - Stack Overflow - https://stackoverflow.com/questions/77855742/how-to-convert-safetensors-model-to-onnx-model
  8. GitHub - justinchuby/onnx-safetensors: Use safetensors with ONNX - https://github.com/justinchuby/onnx-safetensors
  9. Save external data as safetensors - onnx/onnx - Discussion #5461 - https://github.com/onnx/onnx/discussions/5461
  10. GitHub - huggingface/safetensors: Simple, safe way to store and distribute tensors - https://github.com/huggingface/safetensors

distributed training e sharding (FSDP)

tecnicas de sharding de modelo em treino distribuido, FSDP no centro.

  1. Fully Sharded Data Parallel: faster AI training with fewer GPUs Engineering at Meta - https://engineering.fb.com/2021/07/15/open-source/fsdp/
    1. Sharding, Model Parallelism on the IPU with TensorFlow: Sharding and Pipelining - https://docs.graphcore.ai/projects/tf-model-parallelism/en/latest/sharding.html
  2. Sharded Data Parallelism - Amazon SageMaker AI - https://docs.aws.amazon.com/sagemaker/latest/dg/model-parallel-extended-features-pytorch-sharded-data-parallelism.html
  3. [2305.01868] Pre-train and Search: Efficient Embedding Table Sharding with Pre-trained Neural Cost Models - https://arxiv.org/abs/2305.01868
  4. Introducing PyTorch Fully Sharded Data Parallel (FSDP) API, PyTorch - https://pytorch.org/blog/introducing-pytorch-fully-sharded-data-parallel-api/
  5. An Introduction to FSDP (Fully Sharded Data Parallel) for Distributed Training. | by Siddhartha Shrestha | Medium - https://medium.com/@siddharthashrestha/an-introduction-to-fsdp-fully-sharded-data-parallel-for-distributed-training-5e67adfa1712
  6. Decentralized and Distributed Machine Learning Model - https://www.scs.stanford.edu/17au-cs244b/labs/projects/addair.pdf
  7. [2407.19775] Model Agnostic Hybrid Sharding For Heterogeneous Distributed Inference - https://arxiv.org/abs/2407.19775
  8. Fully Sharded Data Parallelism (FSDP) - Edge AI and Vision Alliance - https://www.edge-ai-vision.com/2024/05/fully-sharded-data-parallelism-fsdp/
  9. Getting Started with Fully Sharded Data Parallel (FSDP2), PyTorch Tutorials 2.8.0+cu128 documentation - https://docs.pytorch.org/tutorials/intermediate/FSDP_tutorial.html

model serving, checkpoints e weight storage

vulnerabilidades de serving, formatos de checkpoint e armazenamento de peso em escala. inclui o RAND sobre seguranca de pesos de modelos de fronteira.

  1. Save and load models | TensorFlow Core - https://www.tensorflow.org/tutorials/keras/save_and_load
  2. Architecting scalable checkpoint storage for large-scale ML training on AWS | Amazon Web Services - https://aws.amazon.com/blogs/storage/architecting-scalable-checkpoint-storage-for-large-scale-ml-training-on-aws/
  3. Model Weights File Formats in Machine Learning - https://learnopencv.com/model-weights-file-formats-in-machine-learning/
  4. What is Machine Learning Checkpointing? Deep Learning Models - https://www.deepchecks.com/glossary/machine-learning-checkpointing/
  5. Training checkpoints | TensorFlow Core - https://www.tensorflow.org/guide/checkpoint
  6. What is Machine Learning Checkpointing | Giskard - https://www.giskard.ai/glossary/machine-learning-checkpointing
  7. deep learning - what happens to model weights and how does checkpointing work? - Stack Overflow - https://stackoverflow.com/questions/75368038/what-happens-to-model-weights-and-how-does-checkpointing-work
  8. Securing AI Model Weights: Preventing Theft and Misuse of Frontier Models | RAND - https://www.rand.org/pubs/research_reports/RRA2849-1.html
  9. Saving and Loading Checkpoints, Ray 2.49.0 - https://docs.ray.io/en/latest/train/user-guides/checkpoints.html
  10. Tips and tricks for performing large model checkpointing - https://nebius.com/blog/posts/model-pre-training/large-ml-model-checkpointing-tips

KnockoffNets / ActiveThief (codigo de extracao)

extracao de modelo via active learning sobre dados publicos nao anotados.

  1. ActiveThief: Model Extraction Using Active Learning and Unannotated Public Data | Request PDF - https://www.researchgate.net/publication/341891796_ActiveThief_Model_Extraction_Using_Active_Learning_and_Unannotated_Public_Data

serving infrastructure security e upload de modelos

seguranca de infra de serving de redes neurais, vulnerabilidades e upload de modelos.

  1. How Effective Are Neural Networks for Fixing Security Vulnerabilities | Proceedings of the 32nd ACM SIGSOFT International Symposium on Software Testing and Analysis - https://dl.acm.org/doi/10.1145/3597926.3598135
  2. [2305.18607] How Effective Are Neural Networks for Fixing Security Vulnerabilities - https://arxiv.org/abs/2305.18607
  3. One Step Ahead in Cyber Hide-and-Seek: Automating Malicious Infrastructure Discovery With Graph Neural Networks - https://unit42.paloaltonetworks.com/graph-neural-networks/
  4. A survey on neural networks for (cyber-) security and (cyber-) security of neural networks - ScienceDirect - https://www.sciencedirect.com/science/article/pii/S0925231222007184
  5. Extremely boosted neural network for more accurate multi-stage Cyber attack prediction in cloud computing environment | Journal of Cloud Computing - https://journalofcloudcomputing.springeropen.com/articles/10.1186/s13677-022-00356-9
  6. How Effective Are Neural Networks for Fixing Security Vulnerabilities - https://arxiv.org/html/2305.18607v2
  7. [2011.05976] The Vulnerability of the Neural Networks Against Adversarial Examples in Deep Learning Algorithms - https://arxiv.org/abs/2011.05976
  8. Exploring the Security Vulnerabilities of Neural Networks - https://opendatascience.com/exploring-the-security-vulnerabilities-of-neural-networks/
  9. Machine Learning in Security: Detecting Suspicious Processes Using Recurrent Neural Networks | Splunk - https://www.splunk.com/en_us/blog/security/machine-learning-in-security-detecting-suspicious-processes-using-recurrent-neural-networks.html
  10. Security Assessment of Software Design using Neural - https://arxiv.org/pdf/1303.2017